AION
Research paperMultimodal Models1 source · Oct 7, 2026

Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible.

Key points

  • Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data.
  • But then, do we even need paired examples for cross-modal alignment?
  • Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment.
  • Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs.

Sources (1)

  • [1]Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:12 AM
    Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible.
    Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data.

Extractive summary: sentences quoted from the sources.

Related