The Bitter Lesson of Tool Calling
- Published
- Source
- arXiv
- Paper number
- 849
- Field
- LLMs / NLP
- arXiv ID
- 2608.06370
Key points
- On 11 of 14 models, the Python code method matches or beats the JSON method in accuracy, and the GPT-5.6 family improves by up to 10.6%.
- The code method is overwhelmingly better for chained tool calls, reaching 18.8% higher accuracy than JSON when the chain length is at least 12.
- JSON misses calls when more than 70 tools are invoked at once, while the code method keeps 100% accuracy up to 100 tools.
- Under context rot, the code method remains stable, while JSON drops by 2.3% on average and the filesystem method drops by 32%.
- PTC performance is determined more by model generation time than by model family.
Paper links
External research summaries. These are not HDATF publications or measured product results.