Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given.
ProofPaper ↗
Key points
- We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta).
- Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models).
- This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely.
- Neither better prompting nor direct intervention on the model's internal representations closes this gap.
Sources (1)
- [1]Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance SystemsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:59 PM
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given.
We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta).
Extractive summary: sentences quoted from the sources.