| |
Benchmarks in Leipzig
A group of 49 mathematicians created a benchmark dataset of 100 research-level mathematics questions to evaluate large language models (LLMs), with most work conducted during a 3-day workshop at the Max Planck Institute in Leipzig in April-May 2026. After three evaluation stages using different LLMs, only 2 questions remained completely unsolved, demonstrating significant advances in the mathematical reasoning capabilities of AI systems. The paper includes the full benchmark questions and detailed evaluation statistics across multiple stages of testing.
Read Full Article →
← More Science news