RARE: Retrieval-Augmented Reasoning Modeling
- Published
- Source
- arXiv
- Paper number
- 049
- Field
- Medical AI / RAG
- arXiv ID
- 2503.23513
Key points
- LLMs exhibit knowledge hallucination and insufficient reasoning ability in specialized domains.
- Existing approaches such as RAG mainly supplement knowledge at inference time and do not systematically train domain-specific reasoning during learning.
- Conventional fine-tuning or pretraining for domain knowledge is costly, causes knowledge entanglement, and remains vulnerable to hallucination because of parameter budget limits.
- It externalizes domain knowledge into searchable sources and internalizes domain-specific reasoning patterns inside the LLM.
- It distills high-quality step-by-step reasoning responses that include retrieved knowledge using a teacher model.
- It trains a lightweight student model with a masked-loss objective that injects retrieved knowledge into the prompt, forcing the model to focus on integrating the provided information and reasoning from it.
Paper links
External research summaries. These are not HDATF publications or measured product results.