REFRAG: Rethinking RAG based Decoding
- Published
- Source
- arXiv
- Paper number
- 080
- Field
- RAG / Efficiency
- arXiv ID
- 2509.01092
Key points
- Retrieval-augmented generation (RAG) systems that use large language models (LLMs) suffer from substantial system latency and high memory consumption when processing long retrieved contexts.
- This creates a tradeoff in which enriching knowledge with longer contexts often reduces system efficiency and throughput.
- General-purpose optimization for long-context reasoning in LLMs does not properly handle the RAG-specific block-diagonal attention pattern and redundant computation inherent in RAG contexts.
- A framework using a lightweight encoder compresses RAG contexts into chunk embeddings, greatly reducing the effective input sequence length for a decoder-only LLM.
- The encoder and decoder are aligned through a continual pretraining (CPT) scheme that includes next-paragraph prediction and reconstruction tasks, with curriculum learning used to manage complexity.
- A reinforcement-learning policy dynamically expands important context chunks into full token sequences and compresses less important chunks, enabling flexible compression at arbitrary positions.
Paper links
External research summaries. These are not HDATF publications or measured product results.