SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games.
ProofPaper ↗
Key points
- Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions.
- We study agents' capability of such strategy formation through the communal practice of video game speedrunning.
- Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets.
- These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.
Sources (1)
- [1]SpeedrunBench: Challenging LLM Agents with Video Game SpeedrunningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:06 AM
We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games.
Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions.
Extractive summary: sentences quoted from the sources.