Dr. Zero: Self-Evolving Search Agents without Training Data
- Published
- Source
- arXiv
- Paper number
- 107
- Field
- Agents / Search
- arXiv ID
- 2601.07055
Key points
- A major problem is the need to rely on expensive human-curated training data to develop advanced search and reasoning LLMs.
- Existing data-free self-evolution frameworks have limitations in generating diverse and challenging open-domain questions for multi-hop reasoning.
- Existing reinforcement learning methods for optimizing multi-turn search agents that interact with external tools are computationally inefficient.
- It is a self-evolving proposer-solver framework in which both agents are trained without human data and rely only on external search engines for fact checking.
- Hop-Grouped Relative Policy Optimization (HRPO) is introduced for efficient proposer training, grouping questions by hop complexity and using group-level baselines to reduce compute overhead.
- Difficulty-based rewards are used so the proposer generates diverse, challenging, yet verifiable multi-hop questions, creating a dynamic curriculum.
Paper links
External research summaries. These are not HDATF publications or measured product results.