Skip to contentSkip to stories

Updated

#Reasoning

Sep 8

Sep 8Tue
  1. Mark ChenAI score88

    Mark Chen says OpenAI model helped agents solve Navier-Stokes problem

    AIMark Chen announced that a group of agents produced a solution to the Navier-Stokes Millennium Prize Problem, using an unnamed OpenAI next-generation model. The post says the problem concerns whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and that it had been open for roughly 90 years. The quoted OpenAI post and the attached illustration of inward spiral and axial stretching are cited as context, but the source provides no proof details.

    Why it matters: The post claims an AI-produced proof of a famous open problem, but the source gives no proof details or independent verification, so the claim itself is the main point.

  2. Noam BrownAI score67

    OpenAI shares an AI-generated solution to the Navier-Stokes Millennium Prize Problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem, a question about whether smooth 3D fluid motion can break down that has stayed open for about 90 years. Noam Brown says the result cost millions of dollars, but argues that Astra now scores higher on ARC-AGI for about $20, versus roughly $500,000 for o3 on ARC-AGI 1.

    Why it matters: The post quotes OpenAI's claim about an AI-produced Navier-Stokes solution and adds cost comparisons that show how quickly test-time compute costs are falling.

  3. Noam BrownAI score88

    OpenAI's internal model reportedly solves Navier–Stokes in 88 hours

    AINoam Brown reposted an OpenAI statement that an internal model group reached a Navier–Stokes solution in 88 hours using about 10,000 coordinating AI agents. OpenAI said the model shows a step-function improvement on many benchmarks and that its training is ongoing, with monitoring and isolation safeguards applied throughout. The attached chart compares GPT-6 Astra and the internal model on a curated set of open math problems across test-time compute levels, with the internal model scoring higher at each point.

    Why it matters: The quoted OpenAI post gives concrete figures on an internal model's Navier–Stokes result and on a benchmark comparison, showing how the model performs on open problems.

Sep 7

Sep 7Mon
  1. Tencent HunyuanAI score44

    Tencent Hy4 preview upgraded to cut overthinking and token use

    AITencent Hunyuan says its Hy4 preview has been upgraded to reduce long thinking and over-verification on complex tasks, which users had flagged. The company reports the same task quality with fewer turns and lower input and output tokens, confirmed by benchmark and human evaluation. The upgrade is live for all users, and Tencent says it will keep iterating based on feedback.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 4

Sep 4Fri
  1. John SchulmanAI score34

    Schulman praises metric and dataset for training models to explain behavior

    AIJohn Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline. He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen. That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

Sep 3

Sep 3Thu
  1. Noam BrownAI score50

    OpenAI's Noam Brown Expects GPT-6 Astra to Drive Scientific Discovery

    AINoam Brown, speaking for OpenAI, says he is most excited about GPT-6 Astra's potential for scientific discovery and says OpenAI has not yet pushed the model to its limits on math and science. He looks forward to seeing new scientific breakthroughs built with the model. Background context from a quoted post notes a new OpenAI repo containing a Lean formalization by GPT-6-Astra that proves infinitely many pairs of consecutive primes are at most 186 apart.

Sep 2

Sep 2Wed
  1. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  2. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  3. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  4. Understanding AI (Timothy B. Lee)AI score62

    How Google's RT-2 set the template for today's robotics models

    AIGoogle's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

  5. Google AI DevelopersAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

  6. koray kavukcuogluAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

  7. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

  8. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  9. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  10. Sebastian RaschkaAI score38

    Raschka Says OpenAI Astra's Looped Transformer Is Not a Big Deal

    AISebastian Raschka argues that the looped transformer approach attributed to OpenAI's Astra is a minor architectural tweak, not a major breakthrough. He explains that Nanbeige4.2-3B reuses its 22-layer stack twice, effectively doubling depth without adding weights but roughly doubling compute, and that the idea traces back to the Mixture-of-recursions NeurIPS paper. He adds that layer reuse does not inherently hide chain-of-thought, though it could shift more computation into latent activations.

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

Aug 30

Aug 30Sun
  1. Fireworks AI BlogAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

Aug 29

Aug 29Sat

Aug 28

Aug 28Fri

Aug 27

Aug 27Thu
  1. Leandro von WerraAI score22

    Pollen Robotics unveils Microduck, a $400 open-source RL biped robot

    AIPollen Robotics has unveiled Microduck, a 25 cm open-source biped with 15 actuators and sensors including a camera, speaker, and LiDAR that users can train with reinforcement learning. The robot ships with more than half a dozen pre-trained policies for walking, sitting, roller-skating, and picking up objects with its articulated beak, and costs less than $400.

Aug 26

Aug 26Wed

Aug 21

Aug 21Fri

Aug 19

Aug 19Wed
  1. Jazzyear · InsightsAI score36

    Zhang Yijia Forecasts 2026 AI Trends: Capital Surge and Next-Generation Intelligence Paradigms

    AIZhang Yijia, founder and CEO of Chinese tech think tank 甲子光年, presented a 2026 AI trend report at a Beijing investment conference, arguing China's venture capital has entered a new cycle as financing, investment, and exits rebounded in H1 2026. He said AI absorbed over 70% of global venture investment in H1 2026, with OpenAI and Anthropic together raising $217 billion, and that global AI-related capex is expected to exceed $1 trillion in 2026.

Aug 17

Aug 17Mon
  1. Jason WeiAI score45

    Jason Wei argues tool use cannot replace larger language models

    AIJason Wei now believes a small 1B-parameter "cognitive core" relying on tools is insufficient, because fast, natural recall without tool use matters. He cites speed, knowledge better learned through backpropagation than retrieved from search, and the greater reliability of already-known facts over repeated lookups. Since a 1B model has an information limit, he argues that demanding AI will still need larger models, not just tool access.

  2. Import AIAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 14

Aug 14Fri
  1. Epoch AI · The Epoch BriefAI score42

    Epoch AI lists nine big AI questions its benchmarks aim to answer

    AIEpoch AI outlines nine open questions about AI capabilities, including whether AI can take over full jobs and whether benchmark scores are correlated. The author says Epoch's benchmarking work is built to help answer them, citing examples such as MirrorCode, Remote Labor Index, and the Epoch Capabilities Index (ECI). The post notes that benchmark scores are highly correlated across domains, and that ECI growth trends can help detect whether AI capability progress has accelerated.

Aug 13

Aug 13Thu