AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.
ProofPaper ↗
Key points
- World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction.
- Adapting one such model with a lightweight LoRA fine-tune to a previously unseen robot, a Unitree G1 humanoid with five-fingered BrainCo hands, exposes a video-action asymmetry: the video renders plausible task executions, while the co-generated action is systematically mis-targeted.
- We evaluate closed-loop real-robot trials at three cumulative stages: pre-grasp, grasp, and pick-and-place.
- After one-time embodiment adaptation, self-distillation requires no additional task-specific teleoperation.
Sources (1)
- [1]AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot ActionsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:39 AM
We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.
World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026A Riemannian Geometry for Low-rank Adaptation
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Sep 29, 2026How Diffusion Controller unifies and simplifies AI image generation
- Sep 28, 2026unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
- Sep 28, 2026Notes on NVIDIA Nemotron
- Jun 10, 2026DiffusionGemma: 4x faster text generation