Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison.
ProofPaper ↗
Key points
- Large language models (LLMs) are increasingly used as judges for automated AI evaluation.
- A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear.
- For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity.
- We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
Sources (1)
- [1]Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation EffectsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:53 AM
We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison.
Large language models (LLMs) are increasingly used as judges for automated AI evaluation.
Extractive summary: sentences quoted from the sources.
