LightThinker++: From Reasoning Compression to Memory Management

Published
Source
arXiv
Paper number
143
Field
Reasoning / Efficiency
arXiv ID
2604.03679

Key points

  • LLMs generate many tokens, and the quadratic complexity of attention together with the linear growth of the KV cache creates substantial compute and memory overhead for complex reasoning.
  • Existing context management techniques such as prompt engineering and real-time token pruning often require extensive data construction or introduce significant inference latency.
  • Implicit compression methods are efficient, but they risk irreversible information loss that can create a logical bottleneck in complex reasoning tasks that require detailed intermediate steps.
  • The initial framework, LightThinker, trains the LLM to implicitly compress long reasoning traces into concise semantic representations using special gist tokens and split attention masks.
  • LightThinker++ extends this with explicit, action-level memory management through primitives such as commit, expand, and fold, enabling dynamic preservation of semantic summaries and restoration of raw reasoning detail.
  • It uses a specialized environment-aware trajectory synthesis pipeline to train the model with purposeful memory scheduling and generate expert demonstrations that interleave reasoning with explicit memory operations.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)