Search-o1: Agentic Search-Enhanced Large Reasoning Models
- Published
- Source
- arXiv
- Paper number
- 013
- Field
- Reasoning / Search
- arXiv ID
- 2501.05366
Key points
- As a method, the model pauses generation when it produces the special tokens 'search query start' and 'search query end', retrieves the top k web documents, inserts refined knowledge between the 'search results start' and 'search results end' tokens, and then resumes.
- The key idea is that search is not a one-time preface as in standard RAG, but a recurrent model-controlled action at uncertain points in the reasoning trajectory.
- The Reason-in-Documents stage analyzes retrieved pages in light of the current reasoning context and emits concise final information, preventing verbose or irrelevant pages from derailing the chain of thought.
- In terms of results and implications, the strongest gains appear in knowledge-intensive multi-step settings such as GPQA and multi-hop QA, suggesting that retrieval helps most when the reasoning model can decide both when to search and how to fold evidence back into the proof path.
Paper links
External research summaries. These are not HDATF publications or measured product results.