The Art of Building Verifiers for Computer Use Agents
- Published
- Source
- arXiv
- Paper number
- 139
- Field
- Agents / Evaluation
- arXiv ID
- 2604.06240
Key points
- Reliable verification of task success for Computer Use Agents is difficult because trajectories are long, visually rich, and ambiguous.
- Human annotation of CUA trajectories is tedious and expensive, and it often misses subtle or transient failures.
- Existing automatic verification systems have high false-positive rates because they struggle with token limits, context management, and correctly attributing success or failure in complex environments.
- It introduces Universal Verifier (UV), an evaluation system for CUA trajectories based on four principles of robust verification.
- It uses carefully designed non-overlapping rubrics and two-pass grading to detect hallucinations, and it separates evaluation into process rewards and binary outcome labels.
- UV distinguishes controllable agent failures from uncontrollable environmental factors and uses a divide-and-conquer strategy for multimodal context management with screenshot evidence.
Paper links
External research summaries. These are not HDATF publications or measured product results.