AIME
7stories this week
7last 30 days
7all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceReCal: Calibrating Structured Pruning for On-Policy Distillation RecoveryStructured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
- Oct 8, 2026 · Research paper · 1 sourceSFT-as-Context Mitigates Forgetting in Supervised Fine-TuningWe introduce SFT-as-context, a training-free method in which the parent model uses the SFT model's response as context to answer the query.
- Oct 7, 2026 · Research paper · 1 sourceRSIGym: A Flexible Environment for Recursive Self-ImprovementWe introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS).
- Oct 7, 2026 · Research paper · 1 sourceYANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) MemoryTherefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning.
- Oct 7, 2026 · Research paper · 1 sourceBoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token AggregationWe propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics.
- Oct 7, 2026 · Research paper · 1 sourceFrom Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy DiscoveryTo reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate.
- Oct 7, 2026 · Research paper · 1 sourceCollaborative Reasoning Distillation via Cross-Feedback and Coherent CurationWe propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths.