AION
Research paperLarge Language Models1 source · Oct 7, 2026

Evaluating Rubric Generation with Interventional Transfer

In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them.

Key points

  • Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge.
  • In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks.
  • For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra.
  • Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other.

Sources (1)

  • [1]Evaluating Rubric Generation with Interventional Transfer
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:12 PM
    In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them.
    Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge.

Extractive summary: sentences quoted from the sources.