GSM8K
8stories this week
8last 30 days
11all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceEasy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching missesByte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.
- Oct 8, 2026 · Research paper · 1 sourceAdversarial Cues in Decision Models Used as Judges: The Role of Request PresentationAn answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference.
- Oct 7, 2026 · Research paper · 1 sourceEfficient Best-of-N policy evaluation for inference-time alignmentBest-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model.
- Oct 7, 2026 · Research paper · 1 sourceThe Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance RulesWe treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused.
- Oct 6, 2026 · Research paper · 1 sourceFew Bits, One Law: Toward W2A4KV2We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation.
- Oct 6, 2026 · Research paper · 1 sourceThe Dichotomy Between Pattern Recognition and Step-by-Step ReasoningWe argue that pattern recognition and step-by-step reasoning are two ends of a spectrum.
- Oct 6, 2026 · Research paper · 1 sourceDenoising Hierarchical Representations: Joint Continuous Diffusion for Language ModelingIn this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead.
- Oct 6, 2026 · Research paper · 1 sourceLost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code ChangesTernary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16.
- Sep 5, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.19| Model | Type | PRs | Cookbook |
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- May 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12.post1v0.5.12.post1 is a stability patch on top of v0.5.12.