Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
- Published
- Source
- arXiv
- Paper number
- 073
- Field
- Routing / Efficiency
- arXiv ID
- 2508.12631
Key points
- Balancing high performance, including accuracy and reasoning, with computational efficiency, including cost and latency, is a core challenge in advancing large language models.
- State-of-the-art proprietary LLMs often come with substantial operating costs and resource demands, which limit broad and cost-effective deployment.
- Existing test-time routing methods for LLMs often focus mainly on aggregate performance or add overhead by requiring complex training of additional neural networks.
- The paper develops Avengers-Pro, a test-time routing framework that coordinates diverse LLM ensembles to dynamically select the best model for each query.
- The framework relies on three lightweight unsupervised operations: query embedding, k-means clustering of queries, and per-cluster model-level performance-efficiency scoring.
- During online inference, it embeds incoming queries, assigns them to the nearest cluster, and routes them to the model with the highest aggregate performance-efficiency score, adjustable by the tradeoff parameter α.
Paper links
External research summaries. These are not HDATF publications or measured product results.