Tmax: A simple recipe for terminal agents

Published
Source
arXiv
Paper number
478
Field
LLMs / NLP
arXiv ID
2606.23321

Key points

  • TMAX-15K includes 14,600 RL environment instances, more than 2.5 times the size of the largest existing terminal dataset.
  • A 9B model reaches 27% on Terminal-Bench 2.0, which is the best result among open models under 30B.
  • It uses difficulty control, personas, and validator diversity to synthesize high-quality environments.
  • DPPO, an FP32 LM head, and a large group size stabilize RL for long-horizon agents.
  • It confirms cross-task and cross-harness generalization, including a +5-point gain on SWE-Bench Verified, showing that RL is not just harness fitting.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)