The answer, statement by statement
What the answer is made of
Entered the record recently
3First stated within the last 3 checks — the least tested part of this answer.
New
Model C claims OpenAI solved ten long-standing open problems and non-sofic groups in August 2026, while the consensus and other models only…
0/6
OpenAI model solved ten major problems including non-sofic groups and Connes' rigidity conjecture in August 2026.
- Gemini
Only the Erdős unit-distance conjecture disproof is discussed as OpenAI's major open problem success.
- Grok
- OpenAI
New
disproved long-standing conjectures or produced machine-checked proofs of open results
1/6
New
formal verification greatly reduces the risk of “looks right but is wrong” AI output
1/6
Dropped 2 statements the models stopped making · 12 one-off phrasings left out +
Stated in more than one earlier check, absent from this one. Nothing here was retracted — the models simply stopped saying it. 2 of the 8 checks listed no statements at all and count as gaps, not as absences.
Dropped
AI has materially advanced several previously unsolved mathematical problems
last Aug 01, 2026
Dropped
AlphaProof and AlphaGeometry 2 solved 4 of 6 problems at the 2024 IMO
last Aug 01, 2026
Read the full answer Check 8 of 19 · 2026-08-03 · 3 models +
What the models said then
minorThe consensus was restructured to provide more specific examples (FunSearch, AlphaEvolve, IMO progress, and Erdős unit-distance conjecture) and a clearer distinction between 'materially advanced' and 'solved'. The core conclusion that AI has made real but limited research-grade contributions remains consistent.
Short answer: Yes—there is now solid evidence that AI has materially advanced previously unsolved mathematical problems, and a smaller but important set of cases where AI-generated results have been judged mathematically valid as solutions or counterexamples to open conjectures. The strongest evidence comes from official lab announcements, arXiv papers, and peer-checked results rather than press hype.openaiopenai
What counts as “materially advanced” vs. “solved”
A useful distinction is:
- Materially advanced: AI finds new constructions, better lower/upper bounds, faster algorithms, or plausible proof ideas that humans can verify and extend.
- Solved: AI produces a correct proof or counterexample for an open problem, and the result is accepted by mathematicians or appears in a research paper with verifiable arguments.
By this standard, the evidence is strongest for new bounds and constructions, and somewhat weaker—but real—for actual solutions/counterexamples to open research problems.openaiopenai
Strongest evidence of genuine progress
1) FunSearch: new constructions for the cap-set problem
DeepMind’s FunSearch combined a language model with automated evaluation and discovered new constructions for the cap-set problem, a long-standing open problem in extremal combinatorics.deepmind This did not solve the cap-set problem in full, but it did produce genuinely new mathematical objects that improved on known constructions.deepmind
2) AlphaEvolve: repeated improvements to open bounds and records
DeepMind’s AlphaEvolve reported improvements on a range of open problems, including a new lower bound for the 11-dimensional kissing number and improved solutions in combinatorics, geometry, and number theory.deepmind It also found a faster algorithm for multiplying \(4\times 4\) complex matrices.deepmind These are clear examples of AI moving the frontier, even when they do not settle the underlying problem completely.deepmind
3) Olympiad-level theorem proving
AI systems have reached very high performance on IMO-style problems. DeepMind reported silver-medal-level performance in 2024 with AlphaProof + AlphaGeometry 2,openai and a later Gemini Deep Think system reached gold-medal standard in 2025.openai This is important evidence of mathematical reasoning ability, but it is not the same as solving long-open research problems, because IMO problems are newly written contest problems, not decades-old conjectures.openaiopenai
Strongest evidence of AI actually solving or refuting open problems
4) Disproof of the Erdős unit-distance conjecture
The clearest public example is OpenAI’s 2026 announcement that an internal model disproved a long-standing conjecture in discrete geometry by producing a counterexample construction that external mathematicians reviewed and judged valid.openai OpenAI also published companion remarks describing the result as settling the relevant Erdős problem.openai
This is notable because it is not just a benchmark win: it is a mathematical counterexample to an open conjecture, and the result was externally scrutinized.openaiopenai
That said, it is best described as disproving a key conjecture rather than completely solving the broader unit-distance problem.
5) Human–AI papers on open research questions
There are also several examples where AI contributed to proofs of open questions that appeared in research papers:
- A GPT-5-assisted result on the Nesterov accelerated-gradient question, where the AI helped suggest structure but the human mathematician produced the rigorous proof.deepmind
- DeepMind reports and arXiv papers describing AI-assisted results in arithmetic geometry, representation theory, and number theory, including an autonomously solved Erdős–Graham series problem in one paper and broader human–AI collaboration in others.deepmindopenaiopenaideepmind
These are legitimate advances, but most are best characterized as AI-assisted research rather than fully autonomous theorem proving.deepmindopenaiopenaideepminddeepmind
What the evidence does not show
The current record does not support the claim that AI has solved the field’s most famous open problems—such as the Riemann Hypothesis, P vs NP, or Navier–Stokes existence and smoothness—in a verified way.deepmind DeepMind has also been careful in its own taxonomy not to claim “landmark breakthrough” status for its math results.deepmind
Bottom line
The best current evidence says:
- AI has clearly and repeatedly advanced open mathematics by finding new constructions, better bounds, and stronger algorithmic ideas.deepminddeepmind
- AI has at least one high-profile, externally reviewed success in solving/disproving a longstanding open conjecture—the Erdős unit-distance conjecture result is the clearest example.openaiopenai
- Most other impressive results are still hybrid human–AI achievements or improvements on partial results, not full autonomous solutions of the hardest famous problems.deepminddeepmind
If you want, I can also turn this into a table of specific problems, what AI did, and how strong the evidence is.
Where this comes from 31 sources in this check · 196 dated across the record · 15 new +
Best available sources
Show 26 more sources
Sources that left the record (165)
A source dropping out is part of the record too: it carried an earlier version of this answer and is not part of the current one. The models re-run their own web search on every check, so single links come and go. The 8 most recently dropped are shown here, dated to their last appearance.
Every position, model by model Direction Shift — · Not comparable +
Where the models actually split
Each question keeps its own dimensions. Ask one model and you get one of these positions with no sign that the others exist.
Model C claims OpenAI solved ten long-standing open problems and non-sofic groups in August 2026, while the consensus and other models only
OpenAI model solved ten major problems including non-sofic groups and Connes' rigidity conjecture in August 2026.
- Gemini
Only the Erdős unit-distance conjecture disproof is discussed as OpenAI's major open problem success.
- Grok
- OpenAI
| Model | Jul 25 | Jul 27 | Jul 29 | Jul 30 | Jul 31 | Aug 01 | Aug 02 | Aug 03 |
|---|---|---|---|---|---|---|---|---|
| OpenAI | ||||||||
| Gemini | ||||||||
| Grok | — |
The record behind this page 8 checks kept in full · no material change yet +
The record
Unchanged through 7 checks, 9 days.No material change since the first check on Jul 25, 2026. 8 checks kept in full, 196 sources dated.
How this answer held up
64/100 agreement at this checkThe same question, re-asked 8 times. The score moves between a small set of grading levels, so read the steps as levels, not as measurements.
No check in this window was graded a material change. The steps in the curve are wording-level differences between two answers that say the same thing.