GR2 Technical Report
- Published
- Source
- arXiv
- Paper number
- 536
- Field
- Information Retrieval
- arXiv ID
- 2606.31984
Key points
- Semantic ID mid-training with at least 99% uniqueness lets the LLM directly reason about catalog items.
- Trace distillation from a strong 32B teacher model installs reasoning priors in the 8B student model.
- DAPO-based RL with multiple rewards, format, AUC/NDCG, and LLM-as-judge, optimizes re-ranking.
- It improves industrial traffic by +18.7% in R@1, +7.1% in R@3, and +9.6% in N@3.
- On-Policy Distillation (OPD) recovers about 82% of the 32B teacher gains with a 1.7B model that is only 5% of the size.
- Conditional reward design to prevent reward hacking, such as preserving input order and exploiting position bias, is essential.
Paper links
External research summaries. These are not HDATF publications or measured product results.