Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
We introduce Conditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality.
ProofPaper ↗
Key points
- LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number.
- CAP exposes profile differences hidden by aggregate accuracy: on judgerEva's judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy.
- Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman $\barρ{=}0.87$) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge.
- Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.
Sources (1)
- [1]Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment ConditionsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:46 PM
We introduce Conditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality.
LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number.
Extractive summary: sentences quoted from the sources.