ResearchResearch paperReinforcement Learning · Reasoning & Planning · Robotics & Embodied AI1 source · Oct 6, 2026

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates.

Key points

  • Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences.
  • ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation.
  • Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve.

Sources (1)

  • [1]ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:08 PM
    We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates.
    Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences.

Extractive summary: sentences quoted from the sources.

Related