CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity.
Key points
- Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal.
- It is also susceptible to length and style bias when preference pairs come from different model families.
- To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R.
- On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.
Sources (1)
- [1]CM-DPO: Constraint-Margin Direct Preference Optimization for LLM PlanningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:30 PM
We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity.
Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026llm-openai-decisions 0.1a0
- Oct 6, 2026AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
- Oct 6, 2026Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Sep 29, 2026Introducing GPT-6.1 Sol