AION
Research paperLarge Language Models · Safety & Alignment1 source · Oct 7, 2026

Reliability of LLM Judges for Evaluating Entity Alignment

Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging.

Key points

  • LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied.
  • We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests.
  • A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly.
  • We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA.

Sources (1)

  • [1]Reliability of LLM Judges for Evaluating Entity Alignment
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:55 AM
    Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging.
    LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied.

Extractive summary: sentences quoted from the sources.

Related