AION
Technique

Knowledge distillation

Also known as: distillation, distilled

115stories this week
121last 30 days
125all time

Timeline

  1. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  2. Oct 8, 2026 · Research paper · 1 source
    Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation
    To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
  3. Oct 8, 2026 · Research paper · 2 sources
    One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
    In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
  4. Oct 8, 2026 · Research paper · 2 sources
    ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
    We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
  5. Oct 8, 2026 · Research paper · 1 source
    Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
    Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
  6. Oct 8, 2026 · Research paper · 2 sources
    Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
    Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
  7. Oct 8, 2026 · Research paper · 1 source
    ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration
    We introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction.
  8. Oct 8, 2026 · Research paper · 1 source
    Connected Self Forcing: Beyond Local Learning in Video Autoregression
    To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching.
  9. Oct 8, 2026 · Research paper · 1 source
    Poster: A Preliminary Study of LLM Distillation Inference
    Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers.
  10. Oct 8, 2026 · Research paper · 1 source
    Universal Textual Teaching for LLMs
    We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer.
  11. Oct 8, 2026 · Research paper · 1 source
    Few-Step Generation via Data-Space Iteration
    Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations.
  12. Oct 8, 2026 · Research paper · 1 source
    An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
    PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.
  13. Oct 8, 2026 · Research paper · 1 source
    MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
    In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
  14. Oct 8, 2026 · Research paper · 1 source
    From Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object Interaction
    We propose a hierarchical framework that converts a single-agent HOI policy into a reusable Object-oriented Motion Skill.
  15. Oct 8, 2026 · Research paper · 1 source
    DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
    We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.
  16. Oct 8, 2026 · Research paper · 1 source
    TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction
    Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.
  17. Oct 8, 2026 · Research paper · 1 source
    Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
    We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.
  18. Oct 8, 2026 · Research paper · 1 source
    S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization
    To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.
  19. Oct 8, 2026 · Research paper · 1 source
    SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
    We present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
  20. Oct 8, 2026 · Research paper · 1 source
    ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
    We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
  21. Oct 8, 2026 · Research paper · 1 source
    Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
    We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
  22. Oct 8, 2026 · Research paper · 1 source
    Beyond Sequences: Distilling Structured Decision Memory for LLM Recommendation
    To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence.
  23. Oct 8, 2026 · Research paper · 1 source
    Parametric Trajectory Distillation for Few-Step Video Generation
    We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
  24. Oct 8, 2026 · Research paper · 1 source
    Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
    Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.
  25. Oct 8, 2026 · Research paper · 1 source
    SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
    We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows.
  26. Oct 8, 2026 · Research paper · 1 source
    EchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimation
    We propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.
  27. Oct 8, 2026 · Research paper · 1 source
    Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation
    In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
  28. Oct 8, 2026 · Research paper · 1 source
    Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
    Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.
  29. Oct 8, 2026 · Research paper · 1 source
    PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving
    World models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning.
  30. Oct 8, 2026 · Research paper · 1 source
    Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
    Building on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.

Often appears with