Training AI Scientists to Replicate Research

Published
Source
arXiv
Paper number
905
Field
Machine Learning
arXiv ID
2608.13331

Key points

  • Faraday, a 27B small model, uses a 5-trillion-parameter coding agent as a tool to outperform Claude and GPT-class systems on paper replication, which suggests that it is better to instill scientific judgment in a smaller model than to use a large model directly.
  • For open-ended tasks that cannot be fully verified, it presents a recipe in which an automatic rubric judge aligned with human evaluation is used as the reward for stable reinforcement learning.
  • It achieves a 60 percent win rate and an average gain of 6 percent on held-out AI-for-science tasks, which shows that it transfers scientific reasoning rather than merely memorizing.
  • Faraday does not produce results by blindly coding outputs. Instead, it directly implements the paper's mechanism and scales the experiments down to fit the budget, much like a human scientist.
  • The fact that a weaker model can supervise and direct a stronger model becomes practical evidence relevant to AI safety and oversight research.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)