AION
Research paperLarge Language Models1 source · Oct 8, 2026

Predicting Alignment Generalization with Value Representations

In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values.

Key points

  • LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target.
  • We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values.
  • We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness.
  • Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics.

Sources (1)

  • [1]Predicting Alignment Generalization with Value Representations
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:47 PM
    In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values.
    LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
  2. Oct 8, 2026SuperNav: An Agentic Navigation System for Any Task in Any Scene
  3. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  4. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related