The Bitter Lesson of Tool Calling

Published
Source
arXiv
Paper number
849
Field
LLMs / NLP
arXiv ID
2608.06370

Key points

  • On 11 of 14 models, the Python code method matches or beats the JSON method in accuracy, and the GPT-5.6 family improves by up to 10.6%.
  • The code method is overwhelmingly better for chained tool calls, reaching 18.8% higher accuracy than JSON when the chain length is at least 12.
  • JSON misses calls when more than 70 tools are invoked at once, while the code method keeps 100% accuracy up to 100 tools.
  • Under context rot, the code method remains stable, while JSON drops by 2.3% on average and the filesystem method drops by 32%.
  • PTC performance is determined more by model generation time than by model family.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)