Reliability of LLM Judges for Evaluating Entity Alignment
Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging.
Key points
- LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied.
- We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests.
- A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly.
- We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA.
Sources (1)
- [1]Reliability of LLM Judges for Evaluating Entity AlignmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:55 AM
Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging.
LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied.
Extractive summary: sentences quoted from the sources.