Training AI Co-Scientists Using Rubric Rewards

Published
Source
arXiv
Paper number
103
Field
Scientific AI / RL
arXiv ID
2512.23707

Key points

  • Existing AI applications for science are usually limited to well-defined and executable environments, which reduces their usefulness for abstract and open-ended scientific inquiry.
  • Large language models struggle to generate high-quality long-form research plans because obtaining human feedback in scientific contexts is slow, expensive, and often ambiguous.
  • The main bottleneck is scaling the collection of highly specialized learning data and feedback for open-ended scientific tasks, since manual annotation by human experts is costly and time-consuming.
  • The paper builds an automated pipeline called ResearchPlanGen that extracts research goals, goal-specific grading rubrics, and reference solutions directly from scientific papers to create large and diverse training datasets.
  • It implements a self-graded reinforcement learning framework in which a frozen copy of the initial language model acts as an auto-grader that scores generated plans using privileged goal-specific rubrics and general guidelines.
  • The plan generator is optimized with Group Relative Policy Optimization (GRPO) using the rubric satisfaction score as a direct reward signal, along with a length-control strategy to enforce conciseness.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)