TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests.
ProofPaper ↗
Key points
- Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks.
- To reduce evaluation cost, we also introduce a lightweight variant which audits a random sample of trials in depth and reuses the resulting tests across all trials for the same task.
- Across 6 frontier model backends and 5 benchmarks such as DeepSWE and SWE Marathon, we find that about 34.4% of the model trials currently judged correct violate the task requirements, lowering the overall resolution rate from 50.6% to 33.2%.
- Our results reveal a fundamental limitation of current coding-agent evaluation: as LLMs become better at optimizing against fixed evaluators, those evaluators themselves must become more adaptive.
Sources (1)
- [1]TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator EvolutionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:58 AM
We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests.
Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks.
Extractive summary: sentences quoted from the sources.