AION
Research paperLarge Language Models1 source · Oct 7, 2026

Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged

We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains.

Key points

  • Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation.
  • With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation.
  • In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators.
  • Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation.

Sources (1)

  • [1]Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:17 PM
    We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains.
    Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation.

Extractive summary: sentences quoted from the sources.