CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning

Published
Source
arXiv
Paper number
096
Field
RAG
arXiv ID
2511.18659

Key points

  • Traditional retrieval-augmented generation systems suffer from separate optimization, where the retriever and generator are optimized independently, which creates a mismatch between retrieved information and the generator's actual needs.
  • RAG systems often process long, uncompressed text contexts, which leads to high inference cost, redundant text handling, context overflow in LLMs, and a structural mismatch between embedding-based retrieval and raw-text-based generation.
  • The discrete nature of document selection in retrieval blocks direct gradient flow from the generator back to the retriever, which hinders true end-to-end joint learning and alignment of the retriever with downstream tasks.
  • The Salient Compressor Pretraining (SCP) stage uses synthetic QA pairs and paraphrases to train the compressor to distill documents into concise, semantically rich continuous "memory tokens" that preserve key information.
  • CLaRa integrates retrieval and generation within a single end-to-end differentiable framework by jointly training the query reasoner and generator with a language modeling loss and precomputed compressed document representations.
  • It uses a Straight-Through (ST) estimator to enable gradient flow through discrete top-k document selection, allowing the query reasoner to receive weak supervision from the generator's downstream task objective.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)