Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state.
Key points
- We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities.
- Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items.
- Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year.
- Claims about cross-jurisdiction error need a complete same-state reference set.
Sources (1)
- [1]Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State PolicyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:12 AM
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state.
We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities.
Extractive summary: sentences quoted from the sources.