$α$Transfer: Coefficient Transfer for Efficient Model Merging
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic.
ProofPaper ↗
Key points
- We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes.
- This distributional similarity enables a practical paradigm we call $α$Transfer: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models.
- Experimental results demonstrate a 6$\times$ speedup and 70% memory reduction on vision transformers, and a 20$\times$ speedup and 85% memory reduction on large language models, while maintaining comparable performance.
- Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.
Sources (1)
- [1]$α$Transfer: Coefficient Transfer for Efficient Model MergingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:13 AM
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic.
We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Cloudflare/clef-flash
- Sep 29, 2026microsoft/AesCode-32B
- Sep 29, 2026Language Models for Text Classification: From Bag-of-Words to Jev
- Aug 26, 2026vllm-project/vllm v0.28.0