The answer, statement by statement
What the answer is made of
Restated check after check
2The load-bearing part of this answer. The longest unbroken run on this Topic stands at 1 check.
Settled
found new constructions for the cap-set problem
1/7
Settled
disproved a longstanding conjecture in discrete geometry
1/7
Entered the record recently
1First stated within the last 3 checks — the least tested part of this answer.
New
or other theorem provers to generate machine-checkable proofs.
1/7
Dropped 2 statements the models stopped making · 14 one-off phrasings left out +
Stated in more than one earlier check, absent from this one. Nothing here was retracted — the models simply stopped saying it. 3 of the 10 checks listed no statements at all and count as gaps, not as absences.
Dropped
AI has materially advanced several previously unsolved mathematical problems
last Aug 01, 2026
Dropped
AlphaProof and AlphaGeometry 2 solved 4 of 6 problems at the 2024 IMO
last Aug 01, 2026
Read the full answer Check 10 of 19 · 2026-08-05 · 3 models +
What the models said then
minorThe consensus updated to include a second notable case (an AI-assisted counterexample to the Jacobian conjecture in 2026), added specific details on formal proof campaigns (such as Erdős problems), and refined the categorization of bounds and constructions. The core conclusion remains consistent that AI has advanced mathematical research but not broadly solved the hardest conjectures.
The current evidence says yes, AI has materially advanced some previously unsolved mathematical problems, and in a few cases appears to have solved them outright, but the strongest evidence is still for specific, narrow open problems rather than broad “general mathematical genius.” The most credible public cases are:
1) AI has likely solved at least one prominent open problem
The clearest public example is OpenAI’s report that an internal reasoning model disproved a longstanding conjecture in discrete geometry related to the Erdős unit-distance problem, with a proof checked by external mathematicians and a companion write-up released publicly.openai This is stronger than benchmark performance because it concerns a genuinely open research question, not a contest problem. Independent follow-on mathematical work has already used the construction.deepmind
A second notable case is an AI-assisted counterexample to the Jacobian conjecture in dimension > 2, reported publicly in 2026 and said to have been quickly verified by mathematicians. That would also count as a real resolution of an open problem in the “counterexample” direction, though the public documentation is thinner than for the OpenAI discrete-geometry result.
2) AI has made real progress on open research problems, especially via formal proof and search
Google DeepMind and others have released systems that combine language models with Lean or other theorem provers to generate machine-checkable proofs. One public result reports 9 of 353 open Erdős problems solved in one campaign, plus additional results on OEIS conjectures and related topics.deepmind DeepMind’s own published taxonomy is cautious: it reports results at lower levels of novelty and explicitly does not claim a major “landmark breakthrough” yet.deepmind
This matters because it shows AI can do more than imitate proofs: it can participate in a research loop of proposing, checking, revising, and sometimes completing formal proofs.
3) AI has improved bounds and constructions, even when not fully “solving” the problem
Some of the strongest evidence is in improved constructions rather than final solutions:
- FunSearch found new constructions for the cap-set problem, improving known bounds.deepmind
- AlphaEvolve improved bounds and constructions on several problems, including a new lower bound for a kissing-number problem in 11 dimensions.openai
These are genuine mathematical advances, even though the underlying problems remain open.
4) AI is also helping humans solve research problems faster
There are public examples where AI helped a mathematician or research team close a problem, but the human still framed the question and validated the argument. That is still meaningful progress, but it is better described as AI-assisted research than autonomous AI mathematics.deepmindgithub
Bottom line
The best current evidence supports this summary:
- Confirmed: AI can now produce real, verifiable mathematical advances on open problems.
- Likely/partly confirmed: AI has solved at least one prominent previously open problem and possibly more via counterexamples.
- Not supported yet: the claim that AI routinely performs deep, independent frontier mathematics at the level of the biggest unresolved problems, such as the Millennium Prize Problems.
So the right reading is: AI has moved from contest math into genuine research mathematics, but its strongest successes are still concentrated in problems that are amenable to search, formal verification, counterexample construction, or tight bound improvement.deepmindopenaideepmindopenai
Where this comes from 36 sources in this check · 237 dated across the record · 21 new +
Best available sources
Show 31 more sources
Sources that left the record (201)
A source dropping out is part of the record too: it carried an earlier version of this answer and is not part of the current one. The models re-run their own web search on every check, so single links come and go. The 8 most recently dropped are shown here, dated to their last appearance.
Every position, model by model Direction Shift — · Not comparable +
Where the models actually split
Each question keeps its own dimensions. Ask one model and you get one of these positions with no sign that the others exist.
disproved a longstanding conjecture in discrete geometry
Supported by
- Gemini
- OpenAI
- Grok
or other theorem provers to generate machine-checkable proofs.
Supported by
- Gemini
- OpenAI
- Grok
found new constructions for the cap-set problem
Supported by
- Gemini
- OpenAI
- Grok
| Model | Jul 25 | Jul 27 | Jul 29 | Jul 30 | Jul 31 | Aug 01 | Aug 02 | Aug 03 | Aug 04 | Aug 05 |
|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | ||||||||||
| Gemini | ||||||||||
| Grok | — | — |
The record behind this page 10 checks kept in full · no material change yet +
The record
Unchanged through 9 checks, 11 days.No material change since the first check on Jul 25, 2026. 10 checks kept in full, 237 sources dated.
How this answer held up
84/100 agreement at this checkThe same question, re-asked 10 times. The score moves between a small set of grading levels, so read the steps as levels, not as measurements.
No check in this window was graded a material change. The steps in the curve are wording-level differences between two answers that say the same thing.