Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
Public leaderboards for AI models are read continuously, and attackers can see every published standing.
ProofPaper ↗
Key points
- Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking.
- We introduce the certified corruption budget, a tolerance $\widehat{B}t$ computed after $t$ records and published with each pairwise claim.
- Forged records and records altered once seen require different certificates: the certificate for forgeries fails, with probability approaching one, against an attacker who flips votes it has seen, while one that charges roughly twice as much per record remains valid, with constant bets even against attackers who see the future, and no smaller charge is valid at every level.
- In replays on 1.8 million Chatbot Arena votes, a few hundred rigged votes make standard confidence intervals certify false orderings, while ours stays valid.
Sources (1)
- [1]Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive RiggingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:20 PM
Public leaderboards for AI models are read continuously, and attackers can see every published standing.
Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking.
Extractive summary: sentences quoted from the sources.