Benchmarks in Leipzig

Published
Source
arXiv
Paper number
356
Field
AI / General
arXiv ID
2606.05818

Key points

  • The benchmark contains 100 research-level mathematics problems posed by 49 mathematicians, 35 of whom participated directly in the workshop, spanning 25 subfields including algebraic geometry, combinatorics, and representation theory.
  • It uses a three-stage evaluation pipeline: Stage 1 with 5 models and 1 run each, Stage 2 with 3 models and 20 runs each, and Stage 3 with 2 models and 3 heavy-thinking runs each.
  • The number of unsolved problems drops sharply from 41 to 16 to 2, suggesting that benchmark-style practice has reached the limits of the best-performing models.
  • GPT-5.5 solves 44 problems on a single attempt, the best result, while performance varies dramatically across models, with Grok 4.3 solving only 6 problems.
  • Even repeated runs of the same model show very large performance variance, indicating a deterministic limit in mathematical reasoning.
  • It is a public benchmark based on the ScienceBench platform, and the LLM review process found and removed 3 erroneous problems.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)