ResearchResearch paperEvaluation & Benchmarks · Reasoning & Planning1 source · Oct 7, 2026

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests.

Key points

  • Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks.
  • To reduce evaluation cost, we also introduce a lightweight variant which audits a random sample of trials in depth and reuses the resulting tests across all trials for the same task.
  • Across 6 frontier model backends and 5 benchmarks such as DeepSWE and SWE Marathon, we find that about 34.4% of the model trials currently judged correct violate the task requirements, lowering the overall resolution rate from 50.6% to 33.2%.
  • Our results reveal a fundamental limitation of current coding-agent evaluation: as LLMs become better at optimizing against fixed evaluators, those evaluators themselves must become more adaptive.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related