The answer, statement by statement
What the answer is made of
Entered the record recently
3First stated within the last 3 checks — the least tested part of this answer.
New
Lean-checked proofs of open problems: systems combining LLMs with formal proof assistants have produced machine-verified proofs of some open Erdős problems and related conjectu…
1/8
New
OpenAI announced in May 2026 that an internal model produced a construction violating the conjectured near-linear growth of unit distances among (n) planar points, yielding a low…
1/8
New
The current evidence is that AI has materially advanced several previously open mathematical problems, and in a few cases may have solved them in a credible, externally che…
1/8
Dropped 5 statements the models stopped making · 15 one-off phrasings left out +
Stated in more than one earlier check, absent from this one. Nothing here was retracted — the models simply stopped saying it. 6 of the 14 checks listed no statements at all and count as gaps, not as absences.
Dropped
AI has materially advanced several previously unsolved mathematical problems
last Aug 01, 2026
Dropped
AlphaProof and AlphaGeometry 2 solved 4 of 6 problems at the 2024 IMO
last Aug 01, 2026
Dropped
The specific models or agents credited with solving the planar unit-distance conjecture and other open problems
last Aug 08, 2026
OpenAI’s public claim involved an internal reasoning model disproving the discrete geometry conjecture
- OpenAI
Anthropic's Claude Fable 5 uncovered a counterexample to the Jacobian conjecture
- Gemini
Dropped
disproved a longstanding conjecture in discrete geometry
last Aug 05, 2026
Dropped
found new constructions for the cap-set problem
last Aug 05, 2026
Read the full answer Check 14 of 19 · 2026-08-09 · 3 models +
What the models said then
stableThe update reorganizes the sections, refines specific examples like the Erdős unit-distance conjecture and Lean-checked proofs, and provides slightly more detailed citations and structure, but the core conclusion—that AI has materially advanced mathematics and contributed to open problems but is not yet a general autonomous solver—remains the same.
The current evidence is that AI has materially advanced several previously open mathematical problems, and in a few cases may have solved them in a credible, externally checkable way—but the strongest claims still depend on what counts as “solved” and on how much human verification or cleanup followed the model’s output.
Bottom line
The best-supported examples are:
- A genuine disproof of a long-open conjecture in discrete geometry: OpenAI reported that an internal model found a counterexample to Erdős’s planar unit-distance conjecture, and mathematicians later produced a human-checked version of the argument.openaiopenai
- Lean-checked proofs of open problems: systems combining LLMs with formal proof assistants have produced machine-verified proofs of some open Erdős problems and related conjectures.deepminddeepmind
- Broader but less final research progress: AI systems have improved bounds, constructions, and algorithms in combinatorics and theoretical computer science, even when they did not fully settle the underlying question.deepminddeepmind
What is strongest right now
1) Erdős’s planar unit-distance conjecture
This is the clearest high-profile case. OpenAI announced in May 2026 that an internal model produced a construction violating the conjectured near-linear growth of unit distances among \(n\) planar points, yielding a lower bound of the form \(u(n)\ge n^{1+\delta}\) for some \(\delta>0\).openai A companion mathematical write-up and expert commentary make the argument publicly checkable, and later refinements made the exponent explicit.openai
This is strong evidence that AI can do more than spot patterns: it produced a mathematically substantive counterexample to a long-standing open problem. The remaining caveat is provenance: the final proof is robust, but outsiders cannot fully audit the model’s internal search process.
2) Formal proofs of open problems
A separate, very important line of evidence is formal theorem proving. A preprint on Aristotle/GPT-style proof search reports a Lean formalization of an Erdős problem (problem #728) and describes it as autonomously resolved by an AI system.deepmind Another large-scale effort reported solving a meaningful fraction of formalized open Erdős problems and OEIS conjectures, with Lean proofs and released code.deepmind
These results are especially persuasive because Lean checks the proof mechanically. That said, they are strongest as evidence for AI-assisted formal mathematics rather than for unconstrained “human-style” discovery.
3) Research-level challenge benchmarks
The First Proof benchmark showed that current systems can sometimes produce correct research-level proofs for unpublished problems, including cases judged essentially correct by expert referees.deepmind This is good evidence that AI can now contribute at the level of real research mathematics, not just contest geometry or olympiad-style tasks.
But these benchmarks are not the same as independently discovering a new theorem from scratch: they test performance on problems already selected and formalized by humans, with varying amounts of scaffolding.
What AI has clearly not done yet
There is no solid evidence that AI has broadly solved the hardest open problems in mathematics, such as the major Millennium Prize problems, on its own. Current successes are impressive but still concentrated in areas where:
- the statement is crisp and formalizable;
- the problem can be attacked by search;
- the answer can be checked mechanically or by expert verification;
- existing libraries and formal tools already cover much of the needed ground.
Best overall assessment
AI has definitely materially advanced previously open mathematical problems, and in a few notable cases it has likely resolved them in the mathematical sense. The strongest evidence comes from:
- a counterexample to the Erdős unit-distance conjecture,openaiopenai
- Lean-verified proof search on open Erdős-style problems,deepminddeepmind
- and research-level benchmark successes on unpublished problems.deepmind
So the honest summary is: AI is now a real contributor to frontier mathematics, but not yet a general autonomous solver of open mathematical research.
Where this comes from 26 sources in this check · 290 dated across the record · 9 new +
Best available sources
Show 21 more sources
Sources that left the record (264)
A source dropping out is part of the record too: it carried an earlier version of this answer and is not part of the current one. The models re-run their own web search on every check, so single links come and go. The 8 most recently dropped are shown here, dated to their last appearance.
Every position, model by model Direction Shift 0/100 · Stable +
Where the models actually split
Each question keeps its own dimensions. Ask one model and you get one of these positions with no sign that the others exist.
The current evidence is that AI has materially advanced several previously open mathematical problems, and in a few cases may have solved them in a credible, externally che
Supported by
- OpenAI
- Grok
A genuine disproof of a long-open conjecture in discrete geometry: OpenAI reported that an internal model found a counterexample to Erdős’s planar unit-distance conjecture, and
Supported by
- OpenAI
- Grok
OpenAI reported that an unreleased model family (Astra) generated an AI-driven disproof of the Erdős unit-distance conjecture (specifically constructing a counterexample to Erdős’
- Gemini
Lean-checked proofs of open problems: systems combining LLMs with formal proof assistants have produced machine-verified proofs of some open Erdős problems and related conjectu
Supported by
- OpenAI
- Grok
OpenAI announced in May 2026 that an internal model produced a construction violating the conjectured near-linear growth of unit distances among (n) planar points, yielding a low
Supported by
- OpenAI
- Grok
| Model | Jul 25 | Jul 27 | Jul 29 | Jul 30 | Jul 31 | Aug 01 | Aug 02 | Aug 03 | Aug 04 | Aug 05 | Aug 06 | Aug 07 | Aug 08 | Aug 09 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | ||||||||||||||
| Gemini | ||||||||||||||
| Grok | — | — | — |
The record behind this page 14 checks kept in full · no material change yet +
The record
Unchanged through 13 checks, 15 days.No material change since the first check on Jul 25, 2026. 14 checks kept in full, 290 sources dated.
How this answer held up
82/100 agreement at this checkThe same question, re-asked 14 times. The score moves between a small set of grading levels, so read the steps as levels, not as measurements.
No check in this window was graded a material change. The steps in the curve are wording-level differences between two answers that say the same thing.