SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms.
Key points
- However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints.
- To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective.
- SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior.
- To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits.
Sources (1)
- [1]SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:48 AM
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms.
However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints.
Extractive summary: sentences quoted from the sources.