Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score.
ProofPaper ↗
Key points
- We changed the agent: did it actually get better?
- Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier.
- One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well.
- Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation.
Sources (1)
- [1]Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:13 AM
We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score.
We changed the agent: did it actually get better?
Extractive summary: sentences quoted from the sources.