Dr. Zero: Self-Evolving Search Agents without Training Data

Published
Source
arXiv
Paper number
107
Field
Agents / Search
arXiv ID
2601.07055

Key points

  • A major problem is the need to rely on expensive human-curated training data to develop advanced search and reasoning LLMs.
  • Existing data-free self-evolution frameworks have limitations in generating diverse and challenging open-domain questions for multi-hop reasoning.
  • Existing reinforcement learning methods for optimizing multi-turn search agents that interact with external tools are computationally inefficient.
  • It is a self-evolving proposer-solver framework in which both agents are trained without human data and rely only on external search engines for fact checking.
  • Hop-Grouped Relative Policy Optimization (HRPO) is introduced for efficient proposer training, grouping questions by hop complexity and using group-level baselines to reduce compute overhead.
  • Difficulty-based rewards are used so the proposer generates diverse, challenging, yet verifiable multi-hop questions, creating a dynamic curriculum.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)