Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
Key points
- We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation.
- We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity.
- Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity.
- Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.
Sources (1)
- [1]Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability RetentionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:59 AM
Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Leakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease Research
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
- Oct 6, 2026Are Parameter-Efficient Fine-tuning Methods Really Different?
- Oct 6, 2026MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0