AION
Research paperSafety & Alignment · Large Language Models1 source · Oct 6, 2026

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern.

Key points

  • We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis.
  • We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering.
  • Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP downproj consistently emerging as a prominent safety-sensitive component and oproj providing a smaller contribution.
  • These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.

Sources (1)

  • [1]How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:56 PM
    As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern.
    We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis.

Extractive summary: sentences quoted from the sources.