GR2 Technical Report

Published
Source
arXiv
Paper number
536
Field
Information Retrieval
arXiv ID
2606.31984

Key points

  • Semantic ID mid-training with at least 99% uniqueness lets the LLM directly reason about catalog items.
  • Trace distillation from a strong 32B teacher model installs reasoning priors in the 8B student model.
  • DAPO-based RL with multiple rewards, format, AUC/NDCG, and LLM-as-judge, optimizes re-ranking.
  • It improves industrial traffic by +18.7% in R@1, +7.1% in R@3, and +9.6% in N@3.
  • On-Policy Distillation (OPD) recovers about 82% of the 32B teacher gains with a 1.7B model that is only 5% of the size.
  • Conditional reward design to prevent reward hacking, such as preserving input order and exploiting position bias, is essential.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)