NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
- Published
- Source
- arXiv
- Paper number
- 490
- Field
- LLMs / NLP
- arXiv ID
- 2606.24530
Key points
- It covers 90 carefully selected tasks from 5,500 Nature-family papers across six scientific domains: cell omics, proteins, biological modeling, physical modeling, molecular design, and relational reasoning.
- NatureGym is an automatic pipeline that builds a containerized task environment directly from a paper, which ensures reproducibility.
- Even the strongest agent, Claude Opus 4.7, exceeds the published SOTA only 17.8 percent of the time at g greater than 0.1 and matches it 47.8 percent of the time, which shows that the benchmark relies more on translation than discovery.
- In 45.5 percent of successful paths, the key step was methodological translation rather than scientific invention, meaning conversion into familiar supervised-learning form.
- The main failure causes are wrong method selection at 45.1 percent and insufficient compute budget at 24.4 percent, not lack of task understanding.
- An information firewall hides the original authors' methodology, which forces independent discovery instead of reproduction.
Paper links
External research summaries. These are not HDATF publications or measured product results.