The answer, statement by statement
What the answer is made of
Restated check after check
2The load-bearing part of this answer. The longest unbroken run on this Topic stands at 1 check.
Settled
found new constructions for the cap-set problem
1/7
Settled
disproved a longstanding conjecture in discrete geometry
1/7
Entered the record recently
2First stated within the last 3 checks — the least tested part of this answer.
New
The specific nature and origin of other major open problem breakthroughs besides the Erdős conjecture.
0/7
Model B claims breakthroughs include the Jacobian conjecture via Levent Alpöge and Anthropic's Claude, and AlphaProof Nexus solving 9 out of 353 open Erdős problems.
- Gemini
Model A and Model C focus on DeepMind's Aletheia/Gemini work, FunSearch, and general theorem-proving benchmarks rather than Model B's specific novel claims.
- Grok
- OpenAI
New
or other theorem provers to generate machine-checkable proofs.
1/7
Dropped 2 statements the models stopped making · 14 one-off phrasings left out +
Stated in more than one earlier check, absent from this one. Nothing here was retracted — the models simply stopped saying it. 5 of the 12 checks listed no statements at all and count as gaps, not as absences.
Dropped
AI has materially advanced several previously unsolved mathematical problems
last Aug 01, 2026
Dropped
AlphaProof and AlphaGeometry 2 solved 4 of 6 problems at the 2024 IMO
last Aug 01, 2026
Read the full answer Check 12 of 19 · 2026-08-07 · 3 models +
What the models said then
minorThe consensus was updated to include specific new examples like AlphaEvolve matrix multiplication improvements and IMO gold-medal benchmarks, while reorganizing the primary evidence categories without altering the core conclusion that AI has advanced math without broadly solving it.
Short answer: yes, there is now credible evidence that AI has materially advanced mathematics and has even solved a small number of previously open problems—but not that it has broadly “solved mathematics.” The strongest evidence comes from recent, primary announcements and accompanying formal proofs/benchmarks.
What the evidence shows
1) AI has solved at least one major open problem
The clearest current example is OpenAI’s announcement that an internal reasoning model disproved the Erdős unit-distance conjecture, a long-standing discrete-geometry problem from 1946.openai OpenAI says the proof was externally checked by mathematicians and published with supporting materials, including a proof write-up and formalization.openai That makes this one of the strongest public cases that an AI system has produced a genuine new result on an open mathematical problem.
2) AI has also solved some smaller open Erdős problems
Google DeepMind reported a study on the Erdős problem database in which its agent produced a number of correct results, including a small set that appeared to be novel autonomous solutions rather than literature rediscoveries.deepmindopenai The same work also found that many “open” problems were actually obscure, already-known results, which is an important caution: AI can look impressive while sometimes just recovering hidden literature.openai
3) AI systems have made material advances in constructive mathematics and bounds
DeepMind’s AlphaEvolve produced improved constructions and bounds in several areas, including a rank-48 algorithm for \(4\times4\) matrix multiplication over characteristic-zero fields—an improvement over the long-standing rank-49 baseline.openaideepmind It also improved bounds in combinatorics, uncertainty-type inequalities, packing, and related problems.deepmind These are real mathematical advances even when they do not settle a famous conjecture outright.
4) AI can now solve elite competition problems at or above medal level
DeepMind’s AlphaProof / AlphaGeometry work showed that AI can solve several problems from the International Mathematical Olympiad at silver-medal level, and later DeepMind reported a Gemini Deep Think result at gold-medal level.deepminddeepmind This is strong evidence of advanced reasoning, but these are not previously unsolved research problems in the usual sense.
5) Benchmarks on research-level problems are now more convincing
The First Proof benchmark and related public evaluations show that frontier models can sometimes produce correct proofs or major steps on unpublished research-level problems, though not reliably and not always autonomously.arxiv That suggests the field is moving from competition math toward real research assistance.
What this does not yet show
- AI has not been shown to routinely solve the hardest open problems across mathematics.
- The best results still often depend on formal verification, human selection of problems, careful scaffolding, or post-hoc checking.openaideepmind
- Some claims in the public discussion are overstated; for example, “many landmark breakthroughs” is not supported by DeepMind’s own taxonomy, which still distinguishes these results from top-tier “landmark” advances.deepmind
Bottom line
The current evidence supports this balanced conclusion:
AI has already made genuine, verifiable contributions to unsolved mathematics—ranging from improved bounds and constructions to at least one high-profile disproof of a long-standing conjecture—but it has not yet broadly automated frontier mathematical discovery.openaideepminddeepmind
Key primary sources
- OpenAI, An OpenAI model has disproved a central conjecture in discrete geometry.openai
- Google DeepMind, Accelerating mathematical and scientific discovery with Gemini Deep Think.deepmind
- Google DeepMind, AI solves IMO problems at silver-medal level.deepmind
- Google DeepMind, Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the IMO.deepmind
- OpenAI, Ten advances in mathematics.arxiv
- DeepMind/related paper on Erdős problems and autonomous discovery.deepmindopenai
- DeepMind, AlphaEvolve report and technical paper.openaideepmind
If you want, I can also turn this into a timeline of the most important results from 2023–2026 or a table separating “solved,” “improved,” and “benchmark-only” evidence.
Where this comes from 29 sources in this check · 262 dated across the record · 13 new +
Best available sources
Show 24 more sources
Sources that left the record (233)
A source dropping out is part of the record too: it carried an earlier version of this answer and is not part of the current one. The models re-run their own web search on every check, so single links come and go. The 8 most recently dropped are shown here, dated to their last appearance.
Every position, model by model Direction Shift — · Not comparable +
Where the models actually split
Each question keeps its own dimensions. Ask one model and you get one of these positions with no sign that the others exist.
Model B and Model C describe an OpenAI release called Astra containing ten resolved problems, which is absent from the consensus answer and
Model B and Model C mention an OpenAI Astra framework resolving ten long-standing open problems.
- Grok
- Gemini
The consensus answer and Model A make no mention of an Astra framework.
- OpenAI
| Model | Jul 25 | Jul 27 | Jul 29 | Jul 30 | Jul 31 | Aug 01 | Aug 02 | Aug 03 | Aug 04 | Aug 05 | Aug 06 | Aug 07 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | ||||||||||||
| Gemini | ||||||||||||
| Grok | — | — |
The record behind this page 12 checks kept in full · no material change yet +
The record
Unchanged through 11 checks, 13 days.No material change since the first check on Jul 25, 2026. 12 checks kept in full, 262 sources dated.
How this answer held up
64/100 agreement at this checkThe same question, re-asked 12 times. The score moves between a small set of grading levels, so read the steps as levels, not as measurements.
No check in this window was graded a material change. The steps in the curve are wording-level differences between two answers that say the same thing.