The answer, statement by statement
What the answer is made of
Entered the record recently
2First stated within the last 3 checks — the least tested part of this answer.
New
disproved long-standing conjectures or produced machine-checked proofs of open results
1/6
New
formal verification greatly reduces the risk of “looks right but is wrong” AI output
1/6
Dropped 2 statements the models stopped making · 12 one-off phrasings left out +
Stated in more than one earlier check, absent from this one. Nothing here was retracted — the models simply stopped saying it. 1 of the 7 checks listed no statements at all and count as gaps, not as absences.
Dropped
AI has materially advanced several previously unsolved mathematical problems
last Aug 01, 2026
Dropped
AlphaProof and AlphaGeometry 2 solved 4 of 6 problems at the 2024 IMO
last Aug 01, 2026
Read the full answer Check 7 of 19 · 2026-08-02 · 3 models +
What the models said then
minorThe consensus was updated to include more recent mentions of Lean-verified proofs for open Erdős problems and commutative algebra (such as Rethlas), while maintaining the core conclusion that AI has made real, narrow contributions without broadly solving frontier mathematics.
As of now, the evidence is real but still limited in scope: AI has materially advanced several previously open mathematical problems, and in a few cases has disproved long-standing conjectures or produced machine-checked proofs of open results.deepmindopenaideepmind
What the strongest evidence shows
1) AI has helped resolve some genuine open problems
A notable example is OpenAI’s report that an internal reasoning model produced a proof disproving the planar unit-distance conjecture as it had been commonly understood, yielding an improved lower-bound construction with exponent \(1.014\) rather than \(1+o(1)\).openai
That is a substantive mathematical advance: it is not just solving competition-style exercises, but addressing a published research question.
Another important line of evidence comes from Google DeepMind’s formal-proof work. Their public reports describe systems that, using Lean-based verification, solved a subset of open Erdős problems and other research-level tasks, with results checked in a formal proof assistant.deepmind This matters because formal verification greatly reduces the risk of “looks right but is wrong” AI output.
There are also reports of AI-assisted results in commutative algebra, where a system called Rethlas produced Lean-verified proofs for some open problems.deepmind
2) AI has materially improved known bounds and constructions
Even when a system does not fully “solve” a problem, it can still produce mathematically meaningful progress. Examples include:
- new constructions in additive combinatorics from FunSearch,
- improved algorithms and bounds from AlphaTensor and AlphaEvolve,
- and multiple verified advances in Erdős and OEIS-style problems from DeepMind’s formal-search systems.deepmind
These are best described as material advances, not complete general solutions.
3) The field is moving, but not to “AI has solved mathematics”
The current evidence does not show that AI has broadly solved mathematics, or even that it can routinely crack the hardest famous open problems. The public results are concentrated in areas where:
- the problem can be formalized cleanly,
- correctness can be checked mechanically,
- and the proof search space is amenable to exhaustive or semi-exhaustive exploration.deepminddeepmind
DeepMind’s own classification suggests a cautious reading: its public reports emphasize lower-category results and do not amount to claims of “landmark breakthrough” across mathematics as a whole.deepmind
Bottom line
The best-supported conclusion is:
AI has now crossed from benchmark performance into genuine research-grade mathematical contribution in selected areas, including some open-problem disproofs, verified proofs, and improved bounds/constructions—but it has not generally solved mathematics, and the strongest results are still narrow, formalization-friendly, and often human-checked.deepmindopenaideepmind
If you want, I can also give you a short timeline of the most important public milestones or a table separating “solved,” “disproved,” and “materially advanced” results.
Where this comes from 36 sources in this check · 181 dated across the record · 20 new +
Best available sources
Show 31 more sources
Sources that left the record (145)
A source dropping out is part of the record too: it carried an earlier version of this answer and is not part of the current one. The models re-run their own web search on every check, so single links come and go. The 8 most recently dropped are shown here, dated to their last appearance.
Every position, model by model Direction Shift — · Not comparable +
Where the models actually split
Each question keeps its own dimensions. Ask one model and you get one of these positions with no sign that the others exist.
disproved long-standing conjectures or produced machine-checked proofs of open results
Supported by
- OpenAI
- Gemini
- Grok
OpenAI’s report that an internal reasoning model produced a proof disproving the
Supported by
- OpenAI
- Grok
OpenAI revealed that an internal version of its next-generation reasoning model,
- Gemini
formal verification greatly reduces the risk of “looks right but is wrong” AI output
Supported by
- OpenAI
- Gemini
- Grok
| Model | Jul 25 | Jul 27 | Jul 29 | Jul 30 | Jul 31 | Aug 01 | Aug 02 |
|---|---|---|---|---|---|---|---|
| OpenAI | |||||||
| Gemini | |||||||
| Grok | — |
The record behind this page 7 checks kept in full · no material change yet +
The record
Unchanged through 6 checks, 8 days.No material change since the first check on Jul 25, 2026. 7 checks kept in full, 181 sources dated.
How this answer held up
64/100 agreement at this checkThe same question, re-asked 7 times. The score moves between a small set of grading levels, so read the steps as levels, not as measurements.
No check in this window was graded a material change. The steps in the curve are wording-level differences between two answers that say the same thing.