ResearchResearch paperSafety & Alignment · Large Language Models · Agents & Tool Use1 source · Oct 7, 2026

When AI Finds Hidden Messages, Does It Report?

When an assistant encounters a message for another AI, does it tell its user?

Key points

  • Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions.
  • Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls.
  • Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones.
  • Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user.

Sources (1)

  • [1]When AI Finds Hidden Messages, Does It Report?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:32 AM
    When an assistant encounters a message for another AI, does it tell its user?
    Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions.

Extractive summary: sentences quoted from the sources.

Related