Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
- Published
- Source
- arXiv
- Paper number
- 010
- Field
- RAG / Efficiency
- arXiv ID
- 2412.15605
Key points
- The core idea is to replace online RAG pipelines with a scheme that preloads all relevant documents offline into a long-context LLM and reuses the precomputed KV cache at inference time.
- The method has three stages: knowledge preloading and KV encoding, query-time generation conditioned on the cached state, and cache reset and truncation that removes query-specific tokens while preserving the knowledge cache.
- As a result, on SQuAD 1.0 and HotPotQA in small, medium, and large settings, CAG is reported to match or outperform BM25 and dense-index RAG baselines while eliminating retrieval time. For HotPotQA-Large, generation is reported to take 2.26 seconds instead of 92.08 seconds for dynamic in-context processing.
- In scope and tradeoff, the approach removes the embedding model, vector store, and retriever, which simplifies deployment, but it depends on long-context capacity and is best suited to bounded, relatively stable knowledge bases.
Paper links
External research summaries. These are not HDATF publications or measured product results.