AION
Research paperAgents & Tool Use · Safety & Alignment · Reasoning & Planning1 source · Oct 8, 2026

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.

Key points

  • Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement.
  • A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect.
  • We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset.
  • We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones.

Sources (1)

  • [1]AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:03 AM
    To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
    Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement.

Extractive summary: sentences quoted from the sources.