Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
Key points
- Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time.
- We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it.
- These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable.
- We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.
Sources (2)
- [1]Embodied Turing Machines: Stateful Code for Robot Recursive Self-ImprovementHugging Face Daily Papers · Oct 8, 12:00 AM
We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time.
- [2]Embodied Turing Machines: Stateful Code for Robot Recursive Self-ImprovementarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:29 PM · same content
Extractive summary: sentences quoted from the sources.