ResearchResearch paperLarge Language Models · Interpretability · Multimodal Models1 source · Oct 6, 2026

Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps

LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them?

Key points

  • We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's.
  • Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces.
  • Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts.
  • Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches.

Sources (1)

  • [1]Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:08 PM
    LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them?
    We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's.

Extractive summary: sentences quoted from the sources.

Related