OpenAI releases hundreds of machine-generated math results, drawing verification criticism
Overview
A breakdown of OpenAI's released internal-model math results shows about 73 disproofs and counterexamples, roughly 20% of the total.
The author argues this counters claims that recent math breakthroughs are concentrated in counterexamples because models are only good at brute-force search.
From wh
Check the sources:
Developments
8 developments
- Oct 8, 2:10 PM ET · 1 articleOpenAI releases hundreds of claimed solutions to hard math problems, falling short of mathematicians' advisory standardsTechCrunch · AI: OpenAI's math solutions fall short of the field's standards, mathematicians say
- Oct 7, 6:42 PM ET · 1 articleMathematicians react to hundreds of AI proofs released by OpenAI, noting problems solved in inhuman waysEthan Mollick: Mathematicians react to hundreds of AI-generated proofs released by OpenAI
- Oct 7, 1:16 PM ET · 1 articleGary Marcus criticizes OpenAI's math results disclosure as lacking method detailsMarcus on AI: Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality
- Oct 7, 12:18 PM ET · 1 articleOpenAI's math results spark claims of AGI and Millennium Prize progressDeedy: OpenAI's math results spark claims of AGI and Millennium Prize progress
- Oct 7, 9:50 AM ET · 2 articlesOpenAI's 722 machine-generated math results may split mathematics into two layersExponential View: OpenAI's 722 machine-generated math results may split mathematics into two layers
- Oct 6, 10:44 PM ET · 1 articleMatt Shumer says math is solved, reacting to OpenAI's release of mathematical resultsMatt Shumer: OpenAI releases broad new math results from an internal frontier model
- Oct 6, 9:50 PM ET · 2 articlesOpenAI releases 377 new math results on GitHub alongside a blog postGizmodo · AI: OpenAI Releases 377 Math Results on GitHub Amid Expert Concerns
- Oct 6, 9:03 PM ET · 1 articleAuthor argues 20% counterexample share in OpenAI math results disproves brute-force-search critiquewh: OpenAI's Math Results Are About 20% Disproofs and Counterexamples
Article timeline
Follow the coverage from different perspectives. Times are ET.
- IThome · AI (IT之家)Terence Tao questions OpenAI's 719 AI-generated math proofs
OpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.
- The DecoderMathematicians call for OpenAI boycott after AI-generated proofs flood their field
A group of mathematicians led by Terence Tao has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Tao and other Fields Medalists argue that mass-produced solutions undermine the discipline's focus on conceptual understanding, while Scott Aaronson contrasts this batch release with Anthropic's collaborative approach. The article reports that the internal model tested about 8,000 problems with roughly a five percent success rate.
- TechCrunch · AIOpenAI's math solutions fall short of the field's standards, mathematicians say
OpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.
- Don't Worry About the Vase (Zvi Mowshowitz)AI #189: New Math covers OpenAI's math results and Claude's new pricing
OpenAI reportedly posted solutions to 90 of the top 500 open math problems, using an average of three hours of Pro-level compute per question. Anthropic released Claude Haiku 5.5 at $0.10 input and $0.50 output per million tokens, and the author says Jay Clayton was named AI Czar to head a new taskforce.
- AI Era (新智元)OpenAI publishes 722 AI-generated math manuscripts on GitHub
OpenAI announced on X that it is releasing a series of new mathematical results produced by its internal frontier model. The material is reported as 722 manuscripts posted to GitHub, and the excerpt provides no further detail on the specific results.
- Ethan MollickMathematicians react to hundreds of AI-generated proofs released by OpenAI
Ethan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.
- SiliconANGLE · AIOpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress
OpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.
- Marcus on AIMarcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality
Gary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.
- DeedyOpenAI's math results spark claims of AGI and Millennium Prize progress
Deedy argues LLMs have made substantial progress on four of the seven Millennium Prize problems, including a claimed Navier-Stokes result, conditional on verification. He says OpenAI's results averaged only 3 hours of thinking compute on unreleased models. He concludes that by most definitions of AGI, we have already achieved it.
- The Guardian · AIOpenAI's release of 370 math findings draws expert concern over verification and access
OpenAI published over 370 mathematical results on algebra, theoretical computer science and mathematical logic, drawing concern from mathematicians. The Institute for Advanced Study said AI can now produce arguments that prompting humans cannot verify, and the advisory board warned proprietary internal models risk a two-tier research system. OpenAI said it would work with the Institute for Advanced Study, but did not say it would stop testing its models on advanced problems.
- Exponential ViewOpenAI's 722 machine-generated math results may split mathematics into two layers
OpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.
- Semafor · TechnologyOpenAI's announced math breakthroughs prompt debate over AI's role in proofs
OpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.
- Matt ShumerOpenAI releases broad new math results from an internal frontier model
OpenAI says it is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release them. The results are linked from a GitHub repository at github.com/openai/math.
- Gizmodo · AIOpenAI Releases 377 Math Results on GitHub Amid Expert Concerns
OpenAI released 377 new math results on GitHub, including one paper claiming a proof of the full Birch-Swinnerton-Dyer leading term formula for elliptic curves over the rationals under specific conditions. The results come from the same unreleased internal model that produced its earlier Navier-Stokes result, which conflicts with a September 29 recommendation from the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) to stop testing advanced math problems on proprietary models.
- whOpenAI's Math Results Are About 20% Disproofs and Counterexamples
A breakdown of OpenAI's released internal-model math results shows about 73 disproofs and counterexamples, roughly 20% of the total. The author argues this counters claims that recent math breakthroughs are concentrated in counterexamples because models are only good at brute-force search.
Heat trend
Not enough continuous observations to show a trend yet.