ResearchResearch paperLarge Language Models · Efficiency & Inference1 source · Oct 6, 2026

$α$Transfer: Coefficient Transfer for Efficient Model Merging

Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic.

Key points

  • We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes.
  • This distributional similarity enables a practical paradigm we call $α$Transfer: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models.
  • Experimental results demonstrate a 6$\times$ speedup and 70% memory reduction on vision transformers, and a 20$\times$ speedup and 85% memory reduction on large language models, while maintaining comparable performance.
  • Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.

Sources (1)

  • [1]$α$Transfer: Coefficient Transfer for Efficient Model Merging
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:13 AM
    Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic.
    We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 5, 2026LiquidAI/d1-omni-600M
  2. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  3. Sep 30, 2026Cloudflare/clef-flash
  4. Sep 29, 2026microsoft/AesCode-32B
  5. Sep 29, 2026Language Models for Text Classification: From Bag-of-Words to Jev
  6. Aug 26, 2026vllm-project/vllm v0.28.0

Related