Agents' Last Exam

Published
Source
arXiv
Paper number
345
Field
AI / General
arXiv ID
2606.05405

Key points

  • It is a large benchmark with 55 subdomains, 13 industry clusters, and 1,490 task instances based on the O*NET and SOC 2018 occupational taxonomy.
  • Tasks are collected from projects already completed by working professionals and are evaluated automatically with deterministic scripts and structured rubrics, removing dependence on human raters.
  • The tasks are classified into three tiers by difficulty, Near-Term with 59 tasks, Full-Spectrum with 55, and Last-Exam with 36, to support agent evaluation at different levels.
  • Failure analysis shows that lack of domain knowledge, meaning understanding and approach, accounts for about 75 percent of all failures, so knowledge is the main bottleneck rather than execution ability.
  • Model choice creates about three times larger performance differences than agent harness choice, and using more resources does not always lead to better performance.
  • Even the strongest configuration, Codex plus GPT-5.5, reaches only a 2.6 percent average pass rate, which suggests that the benchmark is still far from saturated.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)