ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning

We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games.

Key points

  • Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions.
  • We study agents' capability of such strategy formation through the communal practice of video game speedrunning.
  • Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets.
  • These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.

Sources (1)

  • [1]SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:06 AM
    We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games.
    Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions.

Extractive summary: sentences quoted from the sources.

Related