Benchmarks in Leipzig
- Published
- Source
- arXiv
- Paper number
- 356
- Field
- AI / General
- arXiv ID
- 2606.05818
Key points
- The benchmark contains 100 research-level mathematics problems posed by 49 mathematicians, 35 of whom participated directly in the workshop, spanning 25 subfields including algebraic geometry, combinatorics, and representation theory.
- It uses a three-stage evaluation pipeline: Stage 1 with 5 models and 1 run each, Stage 2 with 3 models and 20 runs each, and Stage 3 with 2 models and 3 heavy-thinking runs each.
- The number of unsolved problems drops sharply from 41 to 16 to 2, suggesting that benchmark-style practice has reached the limits of the best-performing models.
- GPT-5.5 solves 44 problems on a single attempt, the best result, while performance varies dramatically across models, with Grok 4.3 solving only 6 problems.
- Even repeated runs of the same model show very large performance variance, indicating a deterministic limit in mathematical reasoning.
- It is a public benchmark based on the ScienceBench platform, and the LLM review process found and removed 3 erroneous problems.
Paper links
External research summaries. These are not HDATF publications or measured product results.