End-to-End Context Compression at Scale

Published
Source
arXiv
Paper number
391
Field
LLMs / NLP
arXiv ID
2606.09659

Key points

  • It trains a 0.6B encoder and a 4B decoder based on Qwen3-4B on more than 350B tokens at compression ratios of 1:4, 1:8, and 1:16.
  • A large architecture search optimizes the encoder-decoder design, including pooling, adapter structure, and training stage design.
  • It achieves 8.8x faster TTFT on RULER 4K and 5.2x faster TTFT on LongBench 64K while keeping accuracy ahead.
  • It proposes an agentic harness in which the agent reads the full compressed context at once and uses the EXPAND tool to open only the chunks it needs.
  • It shows near-linear memory scaling up to 1M tokens and is compatible with standard inference engines such as vLLM and SGLang.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)