ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects

We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison.

Key points

  • Large language models (LLMs) are increasingly used as judges for automated AI evaluation.
  • A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear.
  • For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity.
  • We illustrate the approach in an application where AI judges compare two graphical model estimation methods.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related