CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents

Published
Source
arXiv
Paper number
569
Field
Machine Learning
arXiv ID
2607.05378

Key points

  • It integrates context compression, or summarization, into the RL training loop and makes summarization part of a learnable policy.
  • Token-level loss normalization corrects segment length bias, and cross-trajectory GAE assigns credit across compression boundaries.
  • GLM-4.5-Air, a 106B model, reaches 66.8% on SWE-bench Verified, which is a 7.0 point gain, and 24.5% on Terminal-Bench 2.0, which is a 3.1 point gain.
  • GLM-4.7-Flash, a 30B model, reaches 56.0% on SWE-bench Verified, a 5.5 point gain, and 20.2% on Terminal-Bench 2.0, a 6.8 point gain.
  • Including summary training increases summary length, produces richer compressed states, and also increases reasoning tokens, which effectively expands the usable context.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)