CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
- Published
- Source
- arXiv
- Paper number
- 569
- Field
- Machine Learning
- arXiv ID
- 2607.05378
Key points
- It integrates context compression, or summarization, into the RL training loop and makes summarization part of a learnable policy.
- Token-level loss normalization corrects segment length bias, and cross-trajectory GAE assigns credit across compression boundaries.
- GLM-4.5-Air, a 106B model, reaches 66.8% on SWE-bench Verified, which is a 7.0 point gain, and 24.5% on Terminal-Bench 2.0, which is a 3.1 point gain.
- GLM-4.7-Flash, a 30B model, reaches 56.0% on SWE-bench Verified, a 5.5 point gain, and 20.2% on Terminal-Bench 2.0, a 6.8 point gain.
- Including summary training increases summary length, produces richer compressed states, and also increases reasoning tokens, which effectively expands the usable context.
Paper links
External research summaries. These are not HDATF publications or measured product results.