AION
Benchmark

GSM8K

8stories this week
8last 30 days
11all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
    Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
  2. Oct 8, 2026 · Research paper · 1 source
    Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation
    An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference.
  3. Oct 7, 2026 · Research paper · 1 source
    Efficient Best-of-N policy evaluation for inference-time alignment
    Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.
  4. Oct 7, 2026 · Research paper · 1 source
    The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
    We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused.
  5. Oct 6, 2026 · Research paper · 1 source
    Few Bits, One Law: Toward W2A4KV2
    We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation.
  6. Oct 6, 2026 · Research paper · 1 source
    The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
    We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum.
  7. Oct 6, 2026 · Research paper · 1 source
    Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
    In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead.
  8. Oct 6, 2026 · Research paper · 1 source
    Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes
    Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16.
  9. Sep 5, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.19
    | Model | Type | PRs | Cookbook |
  10. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  11. May 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12.post1
    v0.5.12.post1 is a stability patch on top of v0.5.12.

Often appears with