MidTool: Mid-training Data Synthesis for Agentic Tool Use

Published
Source
arXiv
Paper number
976
Field
AI / General
arXiv ID
2608.20314

Key points

  • It open-sourced MidTool-Mix, a 20.3-billion-token mid-training corpus dedicated to tool use that combines web content, PDFs, code, and real API/MCP skills, along with its construction pipeline.
  • Document-based 'grounding' synthesis and real-tool-execution-based 'execution' synthesis play different roles, and both were required for improvements across all eight metrics.
  • Both 4B and 8B models showed consistent gains on three benchmarks, BFCL, tau2-Bench, and MCP Universe, under either SFT or RL post-training.
  • The mid-trained 4B and 8B models outperformed official Qwen3 models on MCP-Universe, showing that dedicated mid-training can be more effective than increasing model size.
  • Execution-trajectory synthesis alone delivered its largest gain on BFCLv3 at +7.9, while grounding synthesis was stronger for transfer to unfamiliar environments such as tau2-Bench and MCP-Universe.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)