Evaluating Rubric Generation with Interventional Transfer
In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them.
Key points
- Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge.
- In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks.
- For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra.
- Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other.
Sources (1)
- [1]Evaluating Rubric Generation with Interventional TransferarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:12 PM
In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them.
Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge.
Extractive summary: sentences quoted from the sources.