AION
Research paperSafety & Alignment1 source · Oct 7, 2026

SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing

Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing.

Key points

  • Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs.
  • We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect.
  • SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries.
  • On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9.

Sources (1)

  • [1]SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:23 PM
    Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing.
    Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs.

Extractive summary: sentences quoted from the sources.