Training AI Scientists to Replicate Research
- Published
- Source
- arXiv
- Paper number
- 905
- Field
- Machine Learning
- arXiv ID
- 2608.13331
Key points
- Faraday, a 27B small model, uses a 5-trillion-parameter coding agent as a tool to outperform Claude and GPT-class systems on paper replication, which suggests that it is better to instill scientific judgment in a smaller model than to use a large model directly.
- For open-ended tasks that cannot be fully verified, it presents a recipe in which an automatic rubric judge aligned with human evaluation is used as the reward for stable reinforcement learning.
- It achieves a 60 percent win rate and an average gain of 6 percent on held-out AI-for-science tasks, which shows that it transfers scientific reasoning rather than merely memorizing.
- Faraday does not produce results by blindly coding outputs. Instead, it directly implements the paper's mechanism and scales the experiments down to fit the budget, much like a human scientist.
- The fact that a weaker model can supervise and direct a stronger model becomes practical evidence relevant to AI safety and oversight research.
Paper links
External research summaries. These are not HDATF publications or measured product results.