ResearchResearch paperRobotics & Embodied AI1 source · Oct 6, 2026

AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions

We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.

Key points

  • World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction.
  • Adapting one such model with a lightweight LoRA fine-tune to a previously unseen robot, a Unitree G1 humanoid with five-fingered BrainCo hands, exposes a video-action asymmetry: the video renders plausible task executions, while the co-generated action is systematically mis-targeted.
  • We evaluate closed-loop real-robot trials at three cumulative stages: pre-grasp, grasp, and pick-and-place.
  • After one-time embodiment adaptation, self-distillation requires no additional task-specific teleoperation.

Sources (1)

  • [1]AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:39 AM
    We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates.
    World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026A Riemannian Geometry for Low-rank Adaptation
  2. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  3. Sep 29, 2026How Diffusion Controller unifies and simplifies AI image generation
  4. Sep 28, 2026unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
  5. Sep 28, 2026Notes on NVIDIA Nemotron
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related