NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

Published
Source
arXiv
Paper number
490
Field
LLMs / NLP
arXiv ID
2606.24530

Key points

  • It covers 90 carefully selected tasks from 5,500 Nature-family papers across six scientific domains: cell omics, proteins, biological modeling, physical modeling, molecular design, and relational reasoning.
  • NatureGym is an automatic pipeline that builds a containerized task environment directly from a paper, which ensures reproducibility.
  • Even the strongest agent, Claude Opus 4.7, exceeds the published SOTA only 17.8 percent of the time at g greater than 0.1 and matches it 47.8 percent of the time, which shows that the benchmark relies more on translation than discovery.
  • In 45.5 percent of successful paths, the key step was methodological translation rather than scientific invention, meaning conversion into familiar supervised-learning form.
  • The main failure causes are wrong method selection at 45.1 percent and insufficient compute budget at 24.4 percent, not lack of task understanding.
  • An information firewall hides the original authors' methodology, which forces independent discovery instead of reproduction.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)