Intern-S2-Preview: Scientific Agentic Foundation Model
- Published
- Source
- arXiv
- Paper number
- 896
- Field
- Machine Learning
- arXiv ID
- 2608.13505
Key points
- We pretrain on rendered scientific documents, including text, figures, and tables, so the model absorbs document structure that is lost in plain text extraction.
- The unified post-training pipeline combines scalable multi-task reinforcement learning with verifiable objectives and black- and white-box agentic RL.
- By freezing the 397B backbone and attaching a separately trained 4B memory decoder, we raise the biology score from 56.92 to 60.32 without hurting general ability.
- The harness-by-task abstraction lets different agent runtimes and tasks share the same rollout, verification, and training protocol.
- It delivers competitive results across scientific, multimodal, agentic, and general benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.