AION
Research paperLarge Language Models1 source · Oct 8, 2026

Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention

Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.

Key points

  • We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation.
  • We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity.
  • Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity.
  • Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.

Sources (1)

  • [1]Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:59 AM
    Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining.
    We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Leakage-Controlled Multimodal Learning for Diagnosis and Progression Prediction in Alzheimer's Disease Research
  2. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  3. Oct 6, 2026The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
  4. Oct 6, 2026Are Parameter-Efficient Fine-tuning Methods Really Different?
  5. Oct 6, 2026MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
  6. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0

Related