Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Published
Source
arXiv
Paper number
1068
Field
AI / General
arXiv ID
2609.04148

Key points

  • Starting from the idea that agent traces already contain all the environmental information, it reconstructed 37.3k reusable execution environments by rewinding file operations and using a repair agent.
  • It generated multiple tasks from each environment, including tasks spanning multiple codebases (breadth expansion) and multi-round tasks with continuing user feedback (depth expansion).
  • By attaching a verifier to each task and using only data that passed, it improved Qwen3.5-27B by 11.9 points on Terminal-Bench 2.1 and 13.8 points on EvoCode-Bench v2 MT@4.
  • With the same data budget, increasing the number of environments was more effective than increasing the number of questions or solutions (53.2→56.0 vs 53.8/53.9).
  • Training the agent to solve tasks again in reconstructed environments was much better than training it to imitate the original traces directly.
  • Data generated from software engineering (SWE) traces also raised the terminal benchmark average from 47.0→50.0, demonstrating cross-domain generalization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)