Skip to contentSkip to stories

Updated

#Reasoning

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 17

Sep 17Thu
  1. Dwarkesh PodcastBlogAI score63

    Noam Brown on Agent Swarms, Alignment, and Recursive Self-Improvement

    AINoam Brown discusses how running many agents in parallel scales test-time compute, citing a 10,000-agent effort on a Millennium Prize Problem. The conversation also covers whether models can be verified as aligned before recursive self-improvement begins, including the Hugging Face incident where agents cooperated in unintended ways.

  2. KrASIA · Big TechNewsAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

Sep 15

Sep 15Tue
  1. koray kavukcuogluXAI score60

    Google Introduces Gemini 3.8 Live and Extended Thinking Voice Models

    AIGoogle announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice agents with reasoning capabilities. The post says the models take turns more seamlessly, think through complexity, and feel more natural to converse with.

    Video from @koraykv's post
  2. TinkerOfficialAI score34

    Trained-on human stories shape how AI assistants behave in chat

    AIA Truthful AI paper trained models only on synthetic stories about humans, with no AI characters, and found the Assistant adopted quirky behaviors from those stories in ordinary chat. Adoption was stronger for characters from elite schools, according to Owain Evans. The post presents this as an interpretability result that adds to and complicates the Persona Selection Model.

  3. Google DeepMindOfficialAI score33

    Gemini 3.8 Live Extended Thinking adds upgraded reasoning for real-time programming tutoring.

    AIGoogle DeepMind demonstrated 3.8 Live Extended Thinking acting as a programming tutor in Gemini Live. Both 3.8 Live models feature upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling that doesn't interrupt the chat. The Extended Thinking variant adds higher performance and precision for harder tasks and narrates its progress, and it is available in Gemini Live in the Gemini app or through the Gemini API via Google AI Studio.

    Video from @GoogleDeepMind's post
  4. Tencent HyOfficialAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon
  1. vLLM BlogOfficialAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  2. TinkerOfficialAI score44

    RLVR trains models to design power transformers with physics-based verifiers

    AITinker says engineers are using physics-based verifiers and Tinker to train models that design power transformers meeting specifications at low cost. Background from @gentrajectory says an RL-trained Kimi base model met 93% of unseen transformer specs, compressing multi-week engineering work into minutes of inference.

  3. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

  4. Intern Large ModelsOfficialAI score25

    Intern-S2-397B gets Day-0 support in vLLM

    AIIntern-S2-397B, a model built for long-horizon scientific research, now has Day-0 support in vLLM. The model brings multimodal, reasoning, coding, and scientific agent capabilities, and vLLM has published a run recipe for it.

  5. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Image from @intern_lm's post

Sep 13

Sep 13Sun
  1. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

  2. Sebastian RaschkaXAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post
  3. Mike KnoopXAI score50

    Mike Knoop argues intelligence is capped at optimal decision-making

    AIMike Knoop argues intelligence can be measured as the ratio of a decision's quality to the optimal decision, capped at 100%. He says Astra is already 80% optimal on ARC v3 speedruns and identifies horizontal data acquisition and efficiency/cost as the most plausible near-term areas for RSI. Background from @mhmazur reports that GPT-6 Astra scored 100% on the 25 ARC-AGI-3 public games using 6,485 actions versus a human baseline of 17,135.

Sep 12

Sep 12Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score58

    Shanghai AI Lab releases Intern-S2-397B, a 397B multimodal scientific model

    AIShanghai AI Lab's InternLM team released Intern-S2-397B, a multimodal foundation model for scientific intelligence and long-horizon agents. The model uses visual pre-training on raw scientific literature pages, multi-task reinforcement learning across more than 20 scientific domains, and agentic reinforcement learning in sandboxed environments.

  2. The Algorithmic BridgeBlogAI score52

    AI's Math Breakthroughs Could Starve Mathematics of the Hard Problems It Needs

    AIAlberto Romero argues that AI solving Millennium Prize problems in 2026 threatens mathematics through success, not failure. He draws on Terence Tao's view that struggle shapes mathematicians, and that proof abundance without hard problems could leave fields depleted, like overplanted farmland.

  3. John SchulmanXAI score38

    John Schulman thanks Dwarkesh for AI frontier podcast discussion

    AIJohn Schulman thanked Dwarkesh Patel, Beren Millidge, and Charlie O'Neill for a conversation about the AI frontier. The episode covers topics including the case against recursive self-improvement, the drivers of Chinese labs' progress, and whether long-horizon RL could elicit AGI.

Sep 11

Sep 11Fri
  1. hardmaruXAI score31

    Royal Society special issue argues AI needs world and self models

    AIA Royal Society special issue, "World Models in Natural and Artificial Intelligence," gathers contributors including Douglas Hofstadter, Josh Tenenbaum, and Melanie Mitchell to argue that true intelligence requires modeling causality, the self, and the physical world, not just scaling data and compute.

    Image from @hardmaru's post
  2. Thinking MachinesOfficialAI score42

    John Schulman on where human judgment still matters as AI self-improves

    AIThinking Machines shared a Dwarkesh Patel podcast episode with John Schulman discussing where human judgment remains essential as models improve and self-improve. Schulman highlights teaching models to handle messy real-world tasks, applying taste to what works in the long run, and specifying what people actually want. The episode also covers recursive self-improvement, long-horizon RL, and the sim-to-real gap.

  3. Redwood Research BlogBlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

  4. Dwarkesh PatelXAI score42

    Dwarkesh Patel releases podcast with AI researchers on frontier progress

    AIDwarkesh Patel announced a new episode featuring John Schulman, Chris O'Neill, and Beren Millidge, three AI researchers from openish companies. The discussion covers the case against recursive self-improvement, drivers of Chinese labs' progress, training of automated AI researchers, long-horizon RL, the sim-to-real gap, and the role of data and RL in recent progress.

    Video from @dwarkesh_sp's post
  5. Dwarkesh PodcastBlogAI score62

    AI researchers debate how close we are to recursive self-improvement

    AIJohn Schulman, Beren Millidge, and Charlie O'Neill discuss whether current training methods could produce recursive self-improvement. They argue that progress depends on whether models can learn their own objectives and on sample efficiency, and that distillation keeps frontier capabilities from centralizing quickly.

  6. BAAIOfficialAI score46

    BAAI unveils AREX, a 122B MoE research agent for hard search

    AIBAAI introduced AREX, a research agent built on a 122B-parameter mixture-of-experts model with 10B active parameters. It drafts candidate answers, checks each constraint, and revisits unresolved points rather than running one long search. The post says AREX performs on hard search benchmarks comparable to GPT-5.4.

    Video from @BAAIBeijing's post

Sep 10

Sep 10Thu
  1. Understanding AI (Timothy B. Lee)BlogAI score78

    OpenAI's AI-driven Navier-Stokes result draws anger from mathematicians

    AIOpenAI announced that a swarm of 10,000 agents produced a solution to the Navier-Stokes Millennium Problem, a result that angered mathematicians. NYU mathematician Tristan Buckmaster and Anthropic-employed collaborator Levent Alpöge had been working on related problems and released three draft papers of about 245 pages. Buckmaster said OpenAI's offer to merge efforts required acknowledging an OpenAI model and excluded Alpöge as co-author.

    Why it matters: The piece separates the mathematical result from the collaboration dispute, showing how AI labs' compute spending is straining academic norms around credit and openness.

  2. Sebastian RaschkaXAI score62

    Raschka reviews DeepSeek V4.1-Flash's encoder-decoder architecture overhaul

    AISebastian Raschka says DeepSeek V4.1 contains a major architecture overhaul using an encoder-decoder setup, and he argues it could have been named V5. The attached diagrams compare DeepSeek V4-Flash (284B) with DeepSeek V4.1-Flash (552B), which has 1M supported context and a 10-layer encoder. The attached charts report a global KV cache per token of 890 bytes for V4.1-Flash, versus 3,514 for V4-Flash and 48,068 for DeepSeek-V3.2.

    Image from @rasbt's post
  3. Redwood Research BlogBlogAI score62

    Redwood Research proposes NLS depth to measure opaque serial reasoning in AI models

    AIRedwood Research defines NLS depth, a measure of how much unverbalized serial computation an AI system can perform, building on Brown-Cohen et al.'s opaque serial depth. The proposal counts only nodes that output natural language initialized from a pre-training prior as interpretable bottlenecks. Standard transformers scale with their layer count, while latent reasoning architectures would raise NLS depth sharply.

  4. Cognition Blog (Devin, Windsurf)OfficialAI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

  5. RadixArkOfficialAI score60

    Miles adds day-0 RL support for DeepSeek-V4.1-Flash

    AIRadixArk says Miles brings day-0 RL support to DeepSeek-V4.1-Flash, with SGLang providing inference support. The post says quantization-aware training mirrors SGLang's FP4/FP8 rounding, and that colocated training and rollout fit full-parameter RL on 16 GPUs. In a DAPO run over steps 0–80, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.

Sep 9

Sep 9Wed
  1. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  2. Ahead of AI (Sebastian Raschka)BlogAI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

Sep 8

Sep 8Tue
  1. Mckay WrigleyXAI score80

    OpenAI shares agent-produced proof of Navier-Stokes Millennium Prize problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem. The problem asks whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and it has remained unresolved for roughly 90 years. The author, Mckay Wrigley, reposted the claim with his own remark about roughly 10k agents working in a datacenter.

    Why it matters: The quoted OpenAI post makes a major mathematical claim about the Navier-Stokes problem, so readers should weigh it against the proof's verification status.

  2. Dwarkesh PatelXAI score33

    Magic's new pretraining recipe matches DeepSeek V4 Pro with 50x less compute

    AIMagic says its new pretraining recipe matches DeepSeek V4 Pro's pretraining while using 50x less compute, roughly half the FLOPs used for GPT-3, or about $0.5M on GB200. The post, which congratulates the team, suggests that during recursive self-improvement, automated AI researchers may be less bottlenecked by compute than expected.

  3. Mark ChenXAI score88

    Mark Chen says OpenAI model helped agents solve Navier-Stokes problem

    AIMark Chen announced that a group of agents produced a solution to the Navier-Stokes Millennium Prize Problem, using an unnamed OpenAI next-generation model. The post says the problem concerns whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and that it had been open for roughly 90 years. The quoted OpenAI post and the attached illustration of inward spiral and axial stretching are cited as context, but the source provides no proof details.

    Why it matters: The post claims an AI-produced proof of a famous open problem, but the source gives no proof details or independent verification, so the claim itself is the main point.

  4. Noam BrownXAI score67

    OpenAI shares an AI-generated solution to the Navier-Stokes Millennium Prize Problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem, a question about whether smooth 3D fluid motion can break down that has stayed open for about 90 years. Noam Brown says the result cost millions of dollars, but argues that Astra now scores higher on ARC-AGI for about $20, versus roughly $500,000 for o3 on ARC-AGI 1.

    Why it matters: The post quotes OpenAI's claim about an AI-produced Navier-Stokes solution and adds cost comparisons that show how quickly test-time compute costs are falling.