When AI Finds Hidden Messages, Does It Report?
When an assistant encounters a message for another AI, does it tell its user?
ProofPaper ↗
Key points
- Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions.
- Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls.
- Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones.
- Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user.
Sources (1)
- [1]When AI Finds Hidden Messages, Does It Report?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:32 AM
When an assistant encounters a message for another AI, does it tell its user?
Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions.
Extractive summary: sentences quoted from the sources.