Tmax: A simple recipe for terminal agents
- Published
- Source
- arXiv
- Paper number
- 478
- Field
- LLMs / NLP
- arXiv ID
- 2606.23321
Key points
- TMAX-15K includes 14,600 RL environment instances, more than 2.5 times the size of the largest existing terminal dataset.
- A 9B model reaches 27% on Terminal-Bench 2.0, which is the best result among open models under 30B.
- It uses difficulty control, personas, and validator diversity to synthesize high-quality environments.
- DPPO, an FP32 LM head, and a large group size stabilize RL for long-horizon agents.
- It confirms cross-task and cross-harness generalization, including a +5-point gain on SWE-Bench Verified, showing that RL is not just harness fitting.
Paper links
External research summaries. These are not HDATF publications or measured product results.