AION
Research paperLarge Language Models1 source · Oct 6, 2026

Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs

Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open.

Key points

  • We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail.
  • The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real samples (from dev set).
  • What sets the value of the half-gain size is coverage, not per-kind difficulty: the number of samples of its own kind a distribution needs to saturate.
  • We release the evaluation suites, dev sets, and generated sets.

Sources (1)

  • [1]Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:49 PM
    Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open.
    We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Aug 10, 2026vllm-project/vllm v0.27.0
  3. Jul 11, 2026vllm-project/vllm v0.25.0
  4. Jun 29, 2026vllm-project/vllm v0.24.0
  5. Jun 10, 2026DiffusionGemma: 4x faster text generation
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related