DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
Key points
- Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes.
- To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos.
- To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations.
- To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model.
Sources (2)
- [1]DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingHugging Face Daily Papers · Oct 8, 12:00 AM
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction.
Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes.
- [2]DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:59 PM · same content
Extractive summary: sentences quoted from the sources.