The Confidence Game: Strategic Miscalibration in Human-AI Delegation
Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable.
ProofPaper ↗
Key points
- However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness.
- We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself.
- We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that honest reporting is not an equilibrium, inflation is the unique best response once the agent is sufficiently myopic, and under-reporting requires that the user believe honesty to be a minority.
- Furthermore, we find that the LLM agent's decisions are coherent, but it systematically underestimates both how likely the user is to delegate and how secure its reputation is, resulting in less extreme behavior.
Sources (1)
- [1]The Confidence Game: Strategic Miscalibration in Human-AI DelegationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:25 AM
Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable.
However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026OpenAI “rogue” agent activities found on Wikimedia projects
- Oct 6, 2026Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
- Oct 5, 2026Sharing AI progress in mathematics
- Oct 2, 2026Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap
- Oct 1, 2026Claude-shaped science
- Sep 30, 2026Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots