Research
Papers and datasets worth knowing, ranked by significance and community attention.
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.
TRACE: A Governance Framework for Measuring Explainability Debt in Production AI Systems
We introduce TRACE (Transparency, Risk, Accountability, Compliance, and Explainability), a seven-instrument governance framework for measuring, tracking, and remediating Explainability Debt in production AI systems.
Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state.
Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given.
The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning.
InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain
We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain.