MidTool: Mid-training Data Synthesis for Agentic Tool Use
- Published
- Source
- arXiv
- Paper number
- 976
- Field
- AI / General
- arXiv ID
- 2608.20314
Key points
- It open-sourced MidTool-Mix, a 20.3-billion-token mid-training corpus dedicated to tool use that combines web content, PDFs, code, and real API/MCP skills, along with its construction pipeline.
- Document-based 'grounding' synthesis and real-tool-execution-based 'execution' synthesis play different roles, and both were required for improvements across all eight metrics.
- Both 4B and 8B models showed consistent gains on three benchmarks, BFCL, tau2-Bench, and MCP Universe, under either SFT or RL post-training.
- The mid-trained 4B and 8B models outperformed official Qwen3 models on MCP-Universe, showing that dedicated mid-training can be more effective than increasing model size.
- Execution-trajectory synthesis alone delivered its largest gain on BFCLv3 at +7.9, while grounding synthesis was stronger for transfer to unfamiliar environments such as tau2-Bench and MCP-Universe.
Paper links
External research summaries. These are not HDATF publications or measured product results.