Research

Papers and datasets worth knowing, ranked by significance and community attention.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness

Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy

When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

TRACE: A Governance Framework for Measuring Explainability Debt in Production AI Systems

We introduce TRACE (Transparency, Risk, Accountability, Compliance, and Explainability), a seven-instrument governance framework for measuring, tracking, and remediating Explainability Debt in production AI systems.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems

Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection

We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain

We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain.

Paper