Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs
Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open.
Key points
- We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail.
- The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real samples (from dev set).
- What sets the value of the half-gain size is coverage, not per-kind difficulty: the number of samples of its own kind a distribution needs to saturate.
- We release the evaluation suites, dev sets, and generated sets.
Sources (1)
- [1]Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe NeedsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:49 PM
Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open.
We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model