PaperGym: Rubric-Centered Evolution for Research-Plan Generation
- Published
- Source
- arXiv
- Paper number
- 1055
- Field
- LLMs / NLP
- arXiv ID
- 2608.31119
Key points
- Using paper structure to separate the sources of questions, drawn from objectives and background, from scoring criteria, drawn from methods and experiments, reduced criterion leakage from the previous 11.9–34.1% to 3.7%.
- Two-stage training that uses the same rubric twice, first as privileged context for OPSD self-teaching and then as a GRPO reward, outperformed SFT, either stage alone, and the reverse order.
- For Qwen3-1.7B/4B/8B, it raised the average across 5 benchmarks by +5.6, +5.0, and +4.8 points, respectively.
- In a controlled comparison with the training recipe fixed, a model trained on PaperGym-20k dominated RubricHub Science (28.2%) with a 58.1% win rate in a three-way comparison, isolating the effect of data quality.
- The trained Qwen3-8B scored 73.48 on ResearchQA, surpassing the much larger Kimi K2.6 at 73.19.
Paper links
External research summaries. These are not HDATF publications or measured product results.