Skip to contentSkip to stories
Updated

#Reasoning

Sep 8

  1. Mckay WrigleyAI score80

    OpenAI shares agent-produced proof of Navier-Stokes Millennium Prize problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem. The problem asks whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and it has remained unresolved for roughly 90 years. The author, Mckay Wrigley, reposted the claim with his own remark about roughly 10k agents working in a datacenter.

    Why it matters: The quoted OpenAI post makes a major mathematical claim about the Navier-Stokes problem, so readers should weigh it against the proof's verification status.

  2. Noam BrownAI score67

    OpenAI shares an AI-generated solution to the Navier-Stokes Millennium Prize Problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem, a question about whether smooth 3D fluid motion can break down that has stayed open for about 90 years. Noam Brown says the result cost millions of dollars, but argues that Astra now scores higher on ARC-AGI for about $20, versus roughly $500,000 for o3 on ARC-AGI 1.

    Why it matters: The post quotes OpenAI's claim about an AI-produced Navier-Stokes solution and adds cost comparisons that show how quickly test-time compute costs are falling.

  3. Noam BrownAI score88

    OpenAI's internal model reportedly solves Navier–Stokes in 88 hours

    AINoam Brown reposted an OpenAI statement that an internal model group reached a Navier–Stokes solution in 88 hours using about 10,000 coordinating AI agents. OpenAI said the model shows a step-function improvement on many benchmarks and that its training is ongoing, with monitoring and isolation safeguards applied throughout. The attached chart compares GPT-6 Astra and the internal model on a curated set of open math problems across test-time compute levels, with the internal model scoring higher at each point.

    Why it matters: The quoted OpenAI post gives concrete figures on an internal model's Navier–Stokes result and on a benchmark comparison, showing how the model performs on open problems.

Apr 24

  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Sep 11, 2024

  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

That’s everything