When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models

Published
Source
arXiv
Paper number
168
Field
Agents / RAG / Reasoning
arXiv ID
2604.26649

Key points

  • Large reasoning models generate long chains of multi-step thought, but they still suffer from factual errors and higher hallucination rates when they need external knowledge.
  • Traditional retrieval-augmented generation systems use a single pre-retrieval step, which fundamentally conflicts with the dynamic and unpredictable knowledge gaps that emerge during an LRM's multi-step reasoning.
  • Existing iterative retrieval methods designed for standard LLMs are not compatible with LRMs because they use the wrong granularity, depend on internal signals unavailable in black-box models, or impose excessive efficiency costs on long reasoning chains.
  • ReaLM-Retrieve, a reasoning-aware retrieval framework, formulates dynamic retrieval as a sequential decision problem and intervenes at the granularity of logical reasoning steps rather than tokens or sentences.
  • It introduces Reasoning Step Uncertainty Score (RSUS), which combines verbalized uncertainty, entity-based entropy, and consistency signals to detect knowledge gaps and retrieval needs in a black-box compatible way.
  • A retrieval intervention policy trained with REINFORCE dynamically decides when to search, constructs relevant queries, and integrates retrieved content efficiently with techniques such as implicit compression, speculative caching, and KV-cache preservation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)