ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
- Published
- Source
- arXiv
- Paper number
- 1091
- Field
- Natural Language Processing
- arXiv ID
- 2609.19134
Key points
- It defined the 'scientific experience bottleneck', where papers and code do not directly become training experience, and proposed infrastructure to solve it.
- It converted 27 scientific codebases into 64 executable environments and 2,812 verified tasks, covering broad fields such as astronomy, climate, and materials.
- The PhAI-IDE model family trained on verified trajectories improved not only on scientific code repair but also on general coding and reasoning benchmarks.
- On the ScienceIDE-Hard evaluation, even the top model (Claude Fable 5.1) reached only about 67%, showing that scientific coding remains a hard area.
- Its aim of an integrated workspace where SFT, reinforcement learning, and evaluation run on the same environment directly resonates with the LabChin architecture.
Paper links
External research summaries. These are not HDATF publications or measured product results.