Cross-Domain Hybrid OPD for Generalizable Search Agents
- Published
- Source
- arXiv
- Paper number
- 805
- Field
- LLMs / NLP
- arXiv ID
- 2608.02101
Key points
- It used two-stage training: first learning autonomous planning and iterative search through agent reinforcement learning, then restoring general capabilities through on-policy distillation from experts in multiple domains.
- Search capabilities were largely preserved after distillation, and several search tasks improved further over the reinforcement-learning stage, demonstrating a balance between specialization and generality.
- Replacing distillation with reinforcement learning on mixed search and general data reduced HYEval3.1 logical reasoning from 46.90% to 36.31% and BBEH Mini from 47.86% to 35.22%.
- Domain-specific teachers and a curriculum beginning with easy problems produced better results on difficult reasoning, coding, and mathematical tasks than a single generalist teacher or training without a curriculum.
- The technical report's evidence focuses on Hunyuan3-based Yuanbao and the presented search and general benchmarks, so generalization to other base models and external operational environments requires additional validation.
Paper links
External research summaries. These are not HDATF publications or measured product results.