Search-o1: Agentic Search-Enhanced Large Reasoning Models

Published
Source
arXiv
Paper number
013
Field
Reasoning / Search
arXiv ID
2501.05366

Key points

  • As a method, the model pauses generation when it produces the special tokens 'search query start' and 'search query end', retrieves the top k web documents, inserts refined knowledge between the 'search results start' and 'search results end' tokens, and then resumes.
  • The key idea is that search is not a one-time preface as in standard RAG, but a recurrent model-controlled action at uncertain points in the reasoning trajectory.
  • The Reason-in-Documents stage analyzes retrieved pages in light of the current reasoning context and emits concise final information, preventing verbose or irrelevant pages from derailing the chain of thought.
  • In terms of results and implications, the strongest gains appear in knowledge-intensive multi-step settings such as GPQA and multi-hop QA, suggesting that retrieval helps most when the reasoning model can decide both when to search and how to fold evidence back into the proof path.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)