CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
- Published
- Source
- arXiv
- Paper number
- 096
- Field
- RAG
- arXiv ID
- 2511.18659
Key points
- Traditional retrieval-augmented generation systems suffer from separate optimization, where the retriever and generator are optimized independently, which creates a mismatch between retrieved information and the generator's actual needs.
- RAG systems often process long, uncompressed text contexts, which leads to high inference cost, redundant text handling, context overflow in LLMs, and a structural mismatch between embedding-based retrieval and raw-text-based generation.
- The discrete nature of document selection in retrieval blocks direct gradient flow from the generator back to the retriever, which hinders true end-to-end joint learning and alignment of the retriever with downstream tasks.
- The Salient Compressor Pretraining (SCP) stage uses synthetic QA pairs and paraphrases to train the compressor to distill documents into concise, semantically rich continuous "memory tokens" that preserve key information.
- CLaRa integrates retrieval and generation within a single end-to-end differentiable framework by jointly training the query reasoner and generator with a language modeling loss and precomputed compressed document representations.
- It uses a Straight-Through (ST) estimator to enable gradient flow through discrete top-k document selection, allowing the query reasoner to receive weak supervision from the generator's downstream task objective.
Paper links
External research summaries. These are not HDATF publications or measured product results.