Training AI Co-Scientists Using Rubric Rewards
- Published
- Source
- arXiv
- Paper number
- 103
- Field
- Scientific AI / RL
- arXiv ID
- 2512.23707
Key points
- Existing AI applications for science are usually limited to well-defined and executable environments, which reduces their usefulness for abstract and open-ended scientific inquiry.
- Large language models struggle to generate high-quality long-form research plans because obtaining human feedback in scientific contexts is slow, expensive, and often ambiguous.
- The main bottleneck is scaling the collection of highly specialized learning data and feedback for open-ended scientific tasks, since manual annotation by human experts is costly and time-consuming.
- The paper builds an automated pipeline called ResearchPlanGen that extracts research goals, goal-specific grading rubrics, and reference solutions directly from scientific papers to create large and diverse training datasets.
- It implements a self-graded reinforcement learning framework in which a frozen copy of the initial language model acts as an auto-grader that scores generated plans using privileged goal-specific rubrics and general guidelines.
- The plan generator is optimized with Group Relative Policy Optimization (GRPO) using the rubric satisfaction score as a direct reward signal, along with a length-control strategy to enforce conciseness.
Paper links
External research summaries. These are not HDATF publications or measured product results.