The Price Reversal Phenomenon: When Cheaper Reasoning Models End Up Costing More

Published
Source
arXiv
Paper number
134
Field
LLMs / Economics
arXiv ID
2603.23971

Key points

  • The method defines real cost as the sum of input-token and output-token costs, includes thinking tokens in billable output, and compares rankings based on list price with rankings based on measured total cost.
  • They audited eight RLMs, including GPT-5.2, Gemini 3 Flash/Pro, Claude Opus/Haiku, Kimi, and MiniMax variants, across 9 datasets and 11,872 queries, using prices from February and March 2026.
  • Among 252 pairwise comparisons, 21.8% showed price reversal. Gemini 3 Flash is 78% cheaper than GPT-5.2 on the published price, but 22% more expensive in total cost, and in a cited extreme case it cost 28 times more than Claude Haiku 4.5 on MMLUPro.
  • RLM procurement should benchmark real workloads and disclose per-request prompt, reasoning, and final-generation cost breakdowns, because even semantic or length-based predictors face stochastic variation in thinking tokens.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)