Retrieval-augmented reasoning with lean language models
- Published
- Source
- arXiv
- Paper number
- 077
- Field
- Medical AI / RAG
- arXiv ID
- 2508.11386
Key points
- State-of-the-art retrieval-augmented generation (RAG) and reasoning systems depend on large external language models, which introduce privacy and security risks in applications that handle sensitive data.
- Deploying and running large frontier language models locally is computationally burdensome for many organizations because of their high resource demands.
- General-purpose language models often lack deep, verifiable expertise in specialized domains, leading to inaccuracy or hallucination in critical applications.
- The researchers used Qwen2.5-Instruct models ranging from 1.5B to 32B parameters as an efficient language-model backbone for reasoning and generation.
- They developed a retrieval system using semantic embeddings, a vector database, and a new strategy that retrieves whole documents and summarizes them in advance to manage context length while preserving relevant information.
- To supervise supervised fine-tuning of the lightweight language models, they generated synthetic domain-specific data with diverse user queries and detailed reasoning trajectories using high-performance models (GPT-4o and DeepSeek-R1).
Paper links
External research summaries. These are not HDATF publications or measured product results.