ResearchResearch paperEvaluation & Benchmarks · Large Language Models · Training & Scaling1 source · Oct 6, 2026

Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions

We introduce Conditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality.

Key points

  • LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number.
  • CAP exposes profile differences hidden by aggregate accuracy: on judgerEva's judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy.
  • Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman $\barρ{=}0.87$) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge.
  • Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.

Sources (1)

  • [1]Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:46 PM
    We introduce Conditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality.
    LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number.

Extractive summary: sentences quoted from the sources.

Related