ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Published
Source
arXiv
Paper number
1091
Field
Natural Language Processing
arXiv ID
2609.19134

Key points

  • It defined the 'scientific experience bottleneck', where papers and code do not directly become training experience, and proposed infrastructure to solve it.
  • It converted 27 scientific codebases into 64 executable environments and 2,812 verified tasks, covering broad fields such as astronomy, climate, and materials.
  • The PhAI-IDE model family trained on verified trajectories improved not only on scientific code repair but also on general coding and reasoning benchmarks.
  • On the ScienceIDE-Hard evaluation, even the top model (Claude Fable 5.1) reached only about 67%, showing that scientific coding remains a hard area.
  • Its aim of an integrated workspace where SFT, reinforcement learning, and evaluation run on the same environment directly resonates with the LabChin architecture.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)