AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

What 30,000 Hours of Ego-centric Video Does Not Teach

We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point.

Key points

  • World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment.
  • We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors.
  • We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains.
  • Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.

Sources (1)

  • [1]What 30,000 Hours of Ego-centric Video Does Not Teach
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:59 PM
    We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point.
    World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment.

Extractive summary: sentences quoted from the sources.