Technique

DPO

Also known as: direct preference optimization

7stories this week
7last 30 days
7all time

Timeline

  1. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  2. Oct 8, 2026 · Research paper · 1 source
    EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
    We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants.
  3. Oct 7, 2026 · Research paper · 1 source
    Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
    Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference.
  4. Oct 7, 2026 · Research paper · 1 source
    From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators
    In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy.
  5. Oct 7, 2026 · Research paper · 1 source
    How to train your model organism
    We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities.
  6. Oct 6, 2026 · Research paper · 1 source
    CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
    We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity.
  7. Oct 6, 2026 · Research paper · 1 source
    AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
    We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.

Often appears with