PaperBench: Evaluating AI's Ability to Replicate AI Research

Published
Source
arXiv
Paper number
046
Field
Benchmarks / Research Agents
arXiv ID
2504.01848

Key points

  • Existing AI agent benchmarks are often too narrow and too simple to evaluate performance on long-horizon, open-ended machine learning R&D tasks.
  • In particular, safety and governance frameworks need rigorous, objective, and scalable measures of an AI model's autonomy and ability to conduct scientific research.
  • Human reproduction of research papers is time-consuming and subjective, which makes it hard to scale for consistent evaluation of AI agents.
  • OpenAI researchers introduced PaperBench, a benchmark of 20 recent ICML papers that requires AI agents to reproduce empirical results from scratch without access to the original codebase.
  • For each paper, a detailed hierarchical rubric developed with the original authors provides binary pass or fail criteria across 8,316 specific contribution items.
  • SimpleJudge, an LLM-based automatic judge validated at an F1 score of 0.83 against human grading benchmark JudgeEval, provides scalable and consistent evaluation of agent submissions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)