Self-Compacting Language Model Agents

Published
Source
arXiv
Paper number
489
Field
LLMs / NLP
arXiv ID
2606.23525

Key points

  • Fixed token-threshold compression can trigger in the middle of reasoning or retrieval and accidentally delete verified facts.
  • The method combines a compression tool that the model can call directly with a lightweight rubric that says to compress when a subtask is solved and suppress compression during reasoning.
  • When the tool is provided without the rubric, usage varies widely across models and is often invoked at the wrong time, which shows that the rubric is essential.
  • On math benchmarks such as IMO-Answerbench and HMMT, it improves by up to 18.1 points over the no-compression baseline.
  • On agentic search tasks such as BrowseComp, it improves by 5 to 9 points while reducing token cost by 30 to 70 percent per question.
  • Adaptive compression is possible at inference time alone, without any fine-tuning or external supervision.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)