A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic
Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models.
ProofPaper ↗
Key points
- However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates.
- Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class.
- Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods.
- Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.
Sources (1)
- [1]A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task ArithmeticarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:49 AM
Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models.
However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates.
Extractive summary: sentences quoted from the sources.