Long Context vs. RAG for LLMs: An Evaluation and Revisits
- Published
- Source
- arXiv
- Paper number
- 012
- Field
- RAG / Long Context
- arXiv ID
- 2501.01880
Key points
- The method uses GPT-4o and strict exact-match checking to filter out questions that can be solved without external context, then evaluates only cases where the provided documents are actually needed.
- The retrieval study compares BM25, Contriever, OpenAI text-embedding-3-small, LlamaIndex-style indexed retrieval, and RAPTOR. RAPTOR is the strongest RAG setup at about 38.5% accuracy.
- As a result, LC reaches about 56.3% exact-match accuracy and outperforms RAG's 49.0% overall, with a clear advantage on factual questions from Wikipedia and stories and on who, where, and which-type questions.
- In tradeoff terms, RAG is still relevant because it works well on fragmented, naturally segmented, dialog-like context and uniquely solves about 10% of evaluation questions that LC fails.
- The evaluation lesson is that synthetic long context assembled from irrelevant passages can bias comparisons because it resembles RAG preprocessing, so context relevance should be treated as a benchmark variable.
Paper links
External research summaries. These are not HDATF publications or measured product results.