Knowledge distillation
Also known as: distillation, distilled
115stories this week
121last 30 days
125all time
Timeline
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 8, 2026 · Research paper · 1 sourceRubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-DistillationTo this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR).
- Oct 8, 2026 · Research paper · 2 sourcesOne Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsIn this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank.
- Oct 8, 2026 · Research paper · 2 sourcesViSkill: Reinforcing VLM Agents with Evolving Visual-Native SkillsWe propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
- Oct 8, 2026 · Research paper · 1 sourceWhich Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-EvolutionSkills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
- Oct 8, 2026 · Research paper · 2 sourcesDistilling Routed 3D Privilege for Spatial Reasoning in Vision-Language ModelsSpatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence.
- Oct 8, 2026 · Research paper · 1 sourceContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image RestorationWe introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction.
- Oct 8, 2026 · Research paper · 1 sourceConnected Self Forcing: Beyond Local Learning in Video AutoregressionTo stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching.
- Oct 8, 2026 · Research paper · 1 sourcePoster: A Preliminary Study of LLM Distillation InferenceUnauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers.
- Oct 8, 2026 · Research paper · 1 sourceUniversal Textual Teaching for LLMsWe introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer.
- Oct 8, 2026 · Research paper · 1 sourceFew-Step Generation via Data-Space IterationFlow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations.
- Oct 8, 2026 · Research paper · 1 sourceAn Interpretable Approach to PDE Solution Discovery via Structural Experience DistillationPDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints.
- Oct 8, 2026 · Research paper · 1 sourceMetaOPD: Meta-Learned Token Weighting for On-Policy DistillationIn this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
- Oct 8, 2026 · Research paper · 1 sourceFrom Solo to Ensemble: A Hierarchical Framework for Composable Multi-Agent Human-Object InteractionWe propose a hierarchical framework that converts a single-agent HOI policy into a reusable Object-oriented Motion Skill.
- Oct 8, 2026 · Research paper · 1 sourceDIAL-OPD: Learning More from Fewer Tokens in On-Policy DistillationWe propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.
- Oct 8, 2026 · Research paper · 1 sourceTAM: Task-Aware Memory Distillation for Efficient Spatiotemporal PredictionKnowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.
- Oct 8, 2026 · Research paper · 1 sourceBeyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible EvidenceWe propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation.
- Oct 8, 2026 · Research paper · 1 sourceS$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-LocalizationTo address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.
- Oct 8, 2026 · Research paper · 1 sourceSDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous DrivingWe present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
- Oct 8, 2026 · Research paper · 1 sourceReTeach: Building a Self-Teacher through Multi-Round Reflection and RetryWe introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification.
- Oct 8, 2026 · Research paper · 1 sourceResidual Advantage: Student-Relative Teacher Guidance for RL with Verifiable RewardsWe propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage.
- Oct 8, 2026 · Research paper · 1 sourceBeyond Sequences: Distilling Structured Decision Memory for LLM RecommendationTo bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence.
- Oct 8, 2026 · Research paper · 1 sourceParametric Trajectory Distillation for Few-Step Video GenerationWe introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path.
- Oct 8, 2026 · Research paper · 1 sourceConditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional TeacherCausal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.
- Oct 8, 2026 · Research paper · 1 sourceSAIL: Scientific Agentic Intelligence via a Science-Aware LoopWe introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows.
- Oct 8, 2026 · Research paper · 1 sourceEchoDiST: Self-distillation-based joint learning for diffusion-conditioned echocardiographic myocardial motion estimationWe propose EchoDiST, a framework for unsupervised echocardiographic myocardial motion estimation that integrates self-distillation-based joint learning with a diffusion-conditioned motion estimation network.
- Oct 8, 2026 · Research paper · 1 sourcePolicy Alignment: New Signals for Membership Auditing in On-Policy DistillationIn this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
- Oct 8, 2026 · Research paper · 1 sourceEnvironmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-DistillationGiven that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce agentic SElf-distilLation with environmental Feedback modeling (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation.
- Oct 8, 2026 · Research paper · 1 sourcePlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous DrivingWorld models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning.
- Oct 8, 2026 · Research paper · 1 sourceRethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text EncoderBuilding on this finding, we propose ComCLIP, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2.