ResearchResearch paperEfficiency & Inference1 source · Oct 7, 2026

Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets

We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces.

Key points

  • Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee.
  • The central observation is validity by abstention: an $(α,δ)$-valid procedure that issues a certificate with probability $P{\rm fire}$ bounds the failure probability of an issued certificate only by $δ/P{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used.
  • We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75π0$ from 0.05 to 0.16, although the floor keeps absolute coverage small.
  • Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $δ$, because abstention absorbs the failures.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
  2. Sep 30, 2026huggingface/transformers v5.18.0: Release 5.18.0
  3. Sep 30, 2026[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU
  4. Sep 29, 2026How to Stop AI Agents From Secretly Collaborating
  5. Jun 16, 2026Unlocking UK house-building with AI-accelerated planning
  6. Jun 16, 2026Securing the future of AI agents

Related