Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study
We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence.
Key points
- Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances.
- A second agent can inspect a completed answer, but the call spends another pool and may provide no substantive finding.
- A boundary-condition scan across both reviewer backends reproduced a previously discovered false success on partial input: four truncation levels passed historically and failed after repair.
- A preregistered shadow study of metadata-only review allocation accrued 25 formal observations before an exact-runtime regression found a third defect: a reviewer exiting nonzero with a well-formed verdict was counted as complete.
Sources (1)
- [1]Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection StudyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:28 PM
We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence.
Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances.
Extractive summary: sentences quoted from the sources.