Agents' Last Exam
- Published
- Source
- arXiv
- Paper number
- 345
- Field
- AI / General
- arXiv ID
- 2606.05405
Key points
- It is a large benchmark with 55 subdomains, 13 industry clusters, and 1,490 task instances based on the O*NET and SOC 2018 occupational taxonomy.
- Tasks are collected from projects already completed by working professionals and are evaluated automatically with deterministic scripts and structured rubrics, removing dependence on human raters.
- The tasks are classified into three tiers by difficulty, Near-Term with 59 tasks, Full-Spectrum with 55, and Last-Exam with 36, to support agent evaluation at different levels.
- Failure analysis shows that lack of domain knowledge, meaning understanding and approach, accounts for about 75 percent of all failures, so knowledge is the main bottleneck rather than execution ability.
- Model choice creates about three times larger performance differences than agent harness choice, and using more resources does not always lead to better performance.
- Even the strongest configuration, Codex plus GPT-5.5, reaches only a 2.6 percent average pass rate, which suggests that the benchmark is still far from saturated.
Paper links
External research summaries. These are not HDATF publications or measured product results.