ResearchResearch paperEfficiency & Inference · Interpretability · Safety & Alignment1 source · Oct 8, 2026

Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception

We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens.

Key points

  • Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people.
  • In one such evaluation, we show that probes can distinguish transcripts containing a model's true hidden goal from other goals with an AUC of up to 99.7%.
  • Our probes also readily detect deception on prominent open-weight models which lie about politically sensitive topics, and about their beliefs when put under pressure.
  • We release our training dataset, dubbed FIBS, to help drive frontier deployment of effective probes, and encourage the community to expand upon it with further examples of deception and sabotage.

Sources (1)

  • [1]Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:58 PM
    We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens.
    Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people.

Extractive summary: sentences quoted from the sources.

Related