PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Published
Source
arXiv
Paper number
1055
Field
LLMs / NLP
arXiv ID
2608.31119

Key points

  • Using paper structure to separate the sources of questions, drawn from objectives and background, from scoring criteria, drawn from methods and experiments, reduced criterion leakage from the previous 11.9–34.1% to 3.7%.
  • Two-stage training that uses the same rubric twice, first as privileged context for OPSD self-teaching and then as a GRPO reward, outperformed SFT, either stage alone, and the reverse order.
  • For Qwen3-1.7B/4B/8B, it raised the average across 5 benchmarks by +5.6, +5.0, and +4.8 points, respectively.
  • In a controlled comparison with the training recipe fixed, a model trained on PaperGym-20k dominated RubricHub Science (28.2%) with a 58.1% win rate in a three-way comparison, isolating the effect of data quality.
  • The trained Qwen3-8B scored 73.48 on ResearchQA, surpassing the much larger Kimi K2.6 at 73.19.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)