SIGMA: Self-Improving Alignment Generalization from a Model Spec
We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings.
Key points
- LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics.
- Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model.
- Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability.
- Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA's task designer agent are crucial for effective self-improvement.
Sources (1)
- [1]SIGMA: Self-Improving Alignment Generalization from a Model SpecarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:09 AM
We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings.
LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 28, 2026openai/openai-python v3.20.0
- Sep 24, 2026openai/openai-python v3.19.2
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation
