The Art of Building Verifiers for Computer Use Agents

Published
Source
arXiv
Paper number
139
Field
Agents / Evaluation
arXiv ID
2604.06240

Key points

  • Reliable verification of task success for Computer Use Agents is difficult because trajectories are long, visually rich, and ambiguous.
  • Human annotation of CUA trajectories is tedious and expensive, and it often misses subtle or transient failures.
  • Existing automatic verification systems have high false-positive rates because they struggle with token limits, context management, and correctly attributing success or failure in complex environments.
  • It introduces Universal Verifier (UV), an evaluation system for CUA trajectories based on four principles of robust verification.
  • It uses carefully designed non-overlapping rubrics and two-pass grading to detect hallucinations, and it separates evaluation into process rewards and binary outcome labels.
  • UV distinguishes controllable agent failures from uncontrollable environmental factors and uses a divide-and-conquer strategy for multimodal context management with screenshot evidence.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)