Skip to contentSkip to stories

Updated

#Reasoning

Oct 8

Oct 8Thu
  1. IThome · AIAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

Oct 7

Oct 7Wed
  1. Semafor · TechnologyAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  2. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

Oct 6

Oct 6Tue
  1. Epoch AIAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    AIEpoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

Sep 30

Sep 30Wed

Sep 27

Sep 27Sun
  1. Tibor BlahoAI score85

    OpenAI releases GPT-6 Sol and Luna as Anthropic launches Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.

    Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.

Sep 25

Sep 25Fri
  1. Kevin WeilAI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

Sep 10

Sep 10Thu
  1. Understanding AI (Timothy B. Lee)AI score78

    OpenAI's AI-driven Navier-Stokes result draws anger from mathematicians

    AIOpenAI announced that a swarm of 10,000 agents produced a solution to the Navier-Stokes Millennium Problem, a result that angered mathematicians. NYU mathematician Tristan Buckmaster and Anthropic-employed collaborator Levent Alpöge had been working on related problems and released three draft papers of about 245 pages. Buckmaster said OpenAI's offer to merge efforts required acknowledging an OpenAI model and excluded Alpöge as co-author.

    Why it matters: The piece separates the mathematical result from the collaboration dispute, showing how AI labs' compute spending is straining academic norms around credit and openness.

Sep 8

Sep 8Tue
  1. Mark ChenAI score88

    Mark Chen says OpenAI model helped agents solve Navier-Stokes problem

    AIMark Chen announced that a group of agents produced a solution to the Navier-Stokes Millennium Prize Problem, using an unnamed OpenAI next-generation model. The post says the problem concerns whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and that it had been open for roughly 90 years. The quoted OpenAI post and the attached illustration of inward spiral and axial stretching are cited as context, but the source provides no proof details.

    Why it matters: The post claims an AI-produced proof of a famous open problem, but the source gives no proof details or independent verification, so the claim itself is the main point.

Jul 20

Jul 20Mon

Jul 19

Jul 19Sun

Jul 5

Jul 5Sun
  1. ARC PrizeAI score47

    ARC Prize Awards First ARC-AGI-3 Milestone Prize to Tufa Labs' Open-Source Agent

    AITufa Labs won the first $37.5K ARC-AGI-3 milestone prize with "The Duck," a small open-source LLM that plays the games by writing and running Python in a live REPL. Reki placed second with a vision-language agent using Gemma-4-31B, and md Boktiar Mahbub Murad placed third with the "forge" framework. The second and final milestone prize ends September 30.

Jun 15

Jun 15Mon
  1. BAAIAI score22

    Turing Award winners Diffie and Barto keynote BAAI Conference on AI security and RL

    AITuring Award winners Whitfield Diffie and Andrew Barto delivered keynotes at the BAAI Conference on AI security and reinforcement learning. Diffie argued that today's feedback-based approach only patches programs after they fail, and that formal methods offer a path to substantially more reliable intended behavior. Barto framed reinforcement learning around control, search, and associative memory, describing its core insight as caching search results rather than searching continuously.

May 21

May 21Thu
  1. Mark ChenAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

Mar 2

Mar 2Mon

Feb 25

Feb 25Wed
  1. Quoc LeAI score53

    Google's Aletheia math agent solves 6 of 10 FirstProof problems

    AIQuoc Le announced that Aletheia, a math research agent, autonomously solved 6 of 10 FirstProof problems, the best result in the inaugural challenge. The post says this exceeds last year's IMO-gold achievement and points to a paper and thread for full details. The accompanying figure shows 10 unmodified problems, 6 candidate solutions per agent, and expert evaluation yielding 6 solved problems on a best-of-2 basis.

Feb 13

Feb 13Fri
  1. Jakub PachockiAI score62

    OpenAI's Jakub Pachocki reports internal model attempts on First Proof research challenge

    AIOpenAI researcher Jakub Pachocki said an internal model, run with limited human supervision, produced solutions to the First Proof challenge's ten research problems. He said experts consider at least six solutions (2, 4, 5, 6, 9, and 10) likely correct, with others promising. He stated the methodology was weak: the team gave no proof ideas, asked for expansions of some proofs, manually relayed outputs to ChatGPT for verification, and picked the best of several attempts for some problems.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

  2. Yi TayAI score38

    Google DeepMind's Gemini team launches new reasoning research group in Singapore

    AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.