AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
Key points
- Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement.
- A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect.
- We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset.
- We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones.
Sources (1)
- [1]AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:03 AM
To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems.
Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement.
Extractive summary: sentences quoted from the sources.