AION
Research paperLarge Language Models · Safety & Alignment1 source · Oct 7, 2026

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms.

Key points

  • However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints.
  • To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective.
  • SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior.
  • To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits.

Sources (1)

  • [1]SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:48 AM
    Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms.
    However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints.

Extractive summary: sentences quoted from the sources.

Related