AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

arXiv cs.AIen

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

arXiv:2610.11050v1 Announce Type: new Abstract: Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect. To identify these errors, a judge needs to examine the trajectory with respect to the user's instruction. To this end, we introduce AgentHorizon, a benchmark of 1,373 compu

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Verktyg
  • Forskning
  • Agenter
  • Företag

Related AI news