AION
Research paperReasoning & Planning · Efficiency & Inference · Robotics & Embodied AI1 source · Oct 8, 2026

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.

Key points

  • The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge.
  • We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims.
  • Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models.
  • By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.

Sources (1)

  • [1]OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:43 AM
    We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature.
    The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge.

Extractive summary: sentences quoted from the sources.