The Price Reversal Phenomenon: When Cheaper Reasoning Models End Up Costing More
- Published
- Source
- arXiv
- Paper number
- 134
- Field
- LLMs / Economics
- arXiv ID
- 2603.23971
Key points
- The method defines real cost as the sum of input-token and output-token costs, includes thinking tokens in billable output, and compares rankings based on list price with rankings based on measured total cost.
- They audited eight RLMs, including GPT-5.2, Gemini 3 Flash/Pro, Claude Opus/Haiku, Kimi, and MiniMax variants, across 9 datasets and 11,872 queries, using prices from February and March 2026.
- Among 252 pairwise comparisons, 21.8% showed price reversal. Gemini 3 Flash is 78% cheaper than GPT-5.2 on the published price, but 22% more expensive in total cost, and in a cited extreme case it cost 28 times more than Claude Haiku 4.5 on MMLUPro.
- RLM procurement should benchmark real workloads and disclose per-request prompt, reasoning, and final-generation cost breakdowns, because even semantic or length-based predictors face stochastic variation in thinking tokens.
Paper links
External research summaries. These are not HDATF publications or measured product results.