CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- Published
- Source
- arXiv
- Paper number
- 471
- Field
- AI / General
- arXiv ID
- 2606.22883
Key points
- The inside-out approach defines tasks from a structured taxonomy first, then grounds them in evidence-based research rather than reusing existing artifacts.
- It uses a three-stage verification process: rubric-gated testing, hint-conditional filtering to remove trivial tasks, and fail-to-pass checks.
- It discards about two-thirds of all candidates, keeping only high-quality tasks to maximize training-signal density.
- Training Qwen3-32B on CLI-Universe-6K, which contains 6,000 trajectories, raises TB 2.0 to 33.4 percent and makes it the best open-source model at or below 32B.
- It improves generalization across benchmarks, with gains of +11.3 on BFCL v4 and +11.6 on VitaBench.
Paper links
External research summaries. These are not HDATF publications or measured product results.