End-to-End Context Compression at Scale
- Published
- Source
- arXiv
- Paper number
- 391
- Field
- LLMs / NLP
- arXiv ID
- 2606.09659
Key points
- It trains a 0.6B encoder and a 4B decoder based on Qwen3-4B on more than 350B tokens at compression ratios of 1:4, 1:8, and 1:16.
- A large architecture search optimizes the encoder-decoder design, including pooling, adapter structure, and training stage design.
- It achieves 8.8x faster TTFT on RULER 4K and 5.2x faster TTFT on LongBench 64K while keeping accuracy ahead.
- It proposes an agentic harness in which the agent reads the full compressed context at once and uses the EXPAND tool to open only the chunks it needs.
- It shows near-linear memory scaling up to 1M tokens and is compatible with standard inference engines such as vLLM and SGLang.
Paper links
External research summaries. These are not HDATF publications or measured product results.