ResearchResearch paperSafety & Alignment1 source · Oct 6, 2026

Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements

We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as Br(N)=arN^alphar, with N a capability proxy; against a budget proportional to N, scaling helps if alphar<1, keeps pace if alphar 1, and accumulates alignment debt if alphar>1.

Key points

  • Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property.
  • We give three operationalizations of burden and distinguish observed, audited and true alignment.
  • We propose a pre-registrable protocol and apply reduced versions of it twice.
  • A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60.

Sources (1)

  • [1]Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:29 PM
    We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r 1, and accumulates alignment debt if alpha_r>1.
    Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  2. Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
  3. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  4. Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
  5. Oct 1, 2026nvidia/PixelUMM
  6. Sep 28, 2026Holo4: powering generalist computer-use agents

Related