Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

Published
Source
arXiv
Paper number
073
Field
Routing / Efficiency
arXiv ID
2508.12631

Key points

  • Balancing high performance, including accuracy and reasoning, with computational efficiency, including cost and latency, is a core challenge in advancing large language models.
  • State-of-the-art proprietary LLMs often come with substantial operating costs and resource demands, which limit broad and cost-effective deployment.
  • Existing test-time routing methods for LLMs often focus mainly on aggregate performance or add overhead by requiring complex training of additional neural networks.
  • The paper develops Avengers-Pro, a test-time routing framework that coordinates diverse LLM ensembles to dynamically select the best model for each query.
  • The framework relies on three lightweight unsupervised operations: query embedding, k-means clustering of queries, and per-cluster model-level performance-efficiency scoring.
  • During online inference, it embeds incoming queries, assigns them to the nearest cluster, and routes them to the model with the highest aggregate performance-efficiency score, adjustable by the tradeoff parameter α.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)