Self-Compacting Language Model Agents
- Published
- Source
- arXiv
- Paper number
- 489
- Field
- LLMs / NLP
- arXiv ID
- 2606.23525
Key points
- Fixed token-threshold compression can trigger in the middle of reasoning or retrieval and accidentally delete verified facts.
- The method combines a compression tool that the model can call directly with a lightweight rubric that says to compress when a subtask is solved and suppress compression during reasoning.
- When the tool is provided without the rubric, usage varies widely across models and is often invoked at the wrong time, which shows that the rubric is essential.
- On math benchmarks such as IMO-Answerbench and HMMT, it improves by up to 18.1 points over the no-compression baseline.
- On agentic search tasks such as BrowseComp, it improves by 5 to 9 points while reducing token cost by 30 to 70 percent per question.
- Adaptive compression is possible at inference time alone, without any fine-tuning or external supervision.
Paper links
External research summaries. These are not HDATF publications or measured product results.