DPO
Also known as: direct preference optimization
7stories this week
7last 30 days
7all time
Timeline
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 8, 2026 · Research paper · 1 sourceEgoVoice: Proactive Spoken Assistance from Egocentric Multimodal StreamsWe introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
- Oct 7, 2026 · Research paper · 1 sourceEnabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion ModelsText-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
- Oct 7, 2026 · Research paper · 1 sourceFrom Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction GeneratorsIn this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy.
- Oct 7, 2026 · Research paper · 1 sourceHow to train your model organismWe re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
- Oct 6, 2026 · Research paper · 1 sourceCM-DPO: Constraint-Margin Direct Preference Optimization for LLM PlanningWe introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity.
- Oct 6, 2026 · Research paper · 1 sourceAutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot ActionsWe propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.