PaperBench: Evaluating AI's Ability to Replicate AI Research
- Published
- Source
- arXiv
- Paper number
- 046
- Field
- Benchmarks / Research Agents
- arXiv ID
- 2504.01848
Key points
- Existing AI agent benchmarks are often too narrow and too simple to evaluate performance on long-horizon, open-ended machine learning R&D tasks.
- In particular, safety and governance frameworks need rigorous, objective, and scalable measures of an AI model's autonomy and ability to conduct scientific research.
- Human reproduction of research papers is time-consuming and subjective, which makes it hard to scale for consistent evaluation of AI agents.
- OpenAI researchers introduced PaperBench, a benchmark of 20 recent ICML papers that requires AI agents to reproduce empirical results from scratch without access to the original codebase.
- For each paper, a detailed hierarchical rubric developed with the original authors provides binary pass or fail criteria across 8,316 specific contribution items.
- SimpleJudge, an LLM-based automatic judge validated at an F1 score of 0.83 against human grading benchmark JudgeEval, provides scalable and consistent evaluation of agent submissions.
Paper links
External research summaries. These are not HDATF publications or measured product results.