AION
Benchmark

AIME

7stories this week
7last 30 days
7all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
    Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
  2. Oct 8, 2026 · Research paper · 1 source
    SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
    We introduce SFT-as-context, a training-free method in which the parent model uses the SFT model's response as context to answer the query.
  3. Oct 7, 2026 · Research paper · 1 source
    RSIGym: A Flexible Environment for Recursive Self-Improvement
    We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS).
  4. Oct 7, 2026 · Research paper · 1 source
    YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory
    Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning.
  5. Oct 7, 2026 · Research paper · 1 source
    BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
    We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
  6. Oct 7, 2026 · Research paper · 1 source
    From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
    To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.
  7. Oct 7, 2026 · Research paper · 1 source
    Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation
    We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths.

Often appears with