Long Context vs. RAG for LLMs: An Evaluation and Revisits

Published
Source
arXiv
Paper number
012
Field
RAG / Long Context
arXiv ID
2501.01880

Key points

  • The method uses GPT-4o and strict exact-match checking to filter out questions that can be solved without external context, then evaluates only cases where the provided documents are actually needed.
  • The retrieval study compares BM25, Contriever, OpenAI text-embedding-3-small, LlamaIndex-style indexed retrieval, and RAPTOR. RAPTOR is the strongest RAG setup at about 38.5% accuracy.
  • As a result, LC reaches about 56.3% exact-match accuracy and outperforms RAG's 49.0% overall, with a clear advantage on factual questions from Wikipedia and stories and on who, where, and which-type questions.
  • In tradeoff terms, RAG is still relevant because it works well on fragmented, naturally segmented, dialog-like context and uniquely solves about 10% of evaluation questions that LC fails.
  • The evaluation lesson is that synthetic long context assembled from irrelevant passages can bias comparisons because it resembles RAG preprocessing, so context relevance should be treated as a benchmark variable.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)