CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents

Published
Source
arXiv
Paper number
471
Field
AI / General
arXiv ID
2606.22883

Key points

  • The inside-out approach defines tasks from a structured taxonomy first, then grounds them in evidence-based research rather than reusing existing artifacts.
  • It uses a three-stage verification process: rubric-gated testing, hint-conditional filtering to remove trivial tasks, and fail-to-pass checks.
  • It discards about two-thirds of all candidates, keeping only high-quality tasks to maximize training-signal density.
  • Training Qwen3-32B on CLI-Universe-6K, which contains 6,000 trajectories, raises TB 2.0 to 33.4 percent and makes it the best open-source model at or below 32B.
  • It improves generalization across benchmarks, with gains of +11.3 on BFCL v4 and +11.6 on VitaBench.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)