Skip to contentSkip to stories

Updated

#Reasoning

Jul 6

Jul 6Mon
  1. Anthropic · YouTubeAI score62

    Anthropic explains how Claude's thoughts split into conscious and automatic levels

    AIAnthropic presents research finding a set of representations in Claude's neural activity that resembles the global workspace theory from neuroscience. The video explains how these representations separate thoughts that are consciously accessible from automatic processing, with a full write-up linked from the source.

    Why it matters: The video explains how Anthropic tested a global workspace analogy inside Claude's neural activity, which bears on how model internals are studied.

Jul 5

Jul 5Sun
  1. ARC PrizeAI score47

    ARC Prize Awards First ARC-AGI-3 Milestone Prize to Tufa Labs' Open-Source Agent

    AITufa Labs won the first $37.5K ARC-AGI-3 milestone prize with "The Duck," a small open-source LLM that plays the games by writing and running Python in a live REPL. Reki placed second with a vision-language agent using Gemma-4-31B, and md Boktiar Mahbub Murad placed third with the "forge" framework. The second and final milestone prize ends September 30.

Jul 1

Jul 1Wed
  1. Mistral AI · new models on Hugging FaceAI score54

    Mistral AI releases Leanstral 1.5, an open-source Lean 4 code agent model

    AIMistral AI released Leanstral 1.5 on Hugging Face as an open-source code agent model for Lean 4 proof assistant tasks. The model uses 119B total parameters with 6.5B activated per token, a 256k context length, and accepts text and image input. The source gives setup paths through Mistral Vibe and a local vLLM server, with recommended settings of temperature 1.0 and reasoning effort set to high for complex prompts. The model is licensed under Apache 2.0.

Jun 30

Jun 30Tue
  1. Jim FanAI score60

    ASPIRE lets robots build an evolving skills library that transfers across tasks

    AIJim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.

Jun 29

Jun 29Mon
  1. Meta AI BlogAI score68

    Meta's Brain2Qwerty v2 decodes sentences from non-invasive brain recordings

    AIMeta released Brain2Qwerty v2, an end-to-end deep learning pipeline that decodes sentences in real time from non-invasive brain recordings. The model reached 61% word accuracy across participants, compared with 8% for other non-invasive methods, and 78% for the best participant. Meta also released the v1 and v2 training code, and partner BCBL released the v1 dataset.

    Why it matters: The source reports word accuracy and data-scaling results for non-invasive decoding, offering a benchmark against surgical brain-computer interfaces and prior non-invasive methods.

Jun 26

Jun 26Fri

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    AIOpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    Why it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 17

Jun 17Wed

Jun 16

Jun 16Tue
  1. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. BAAIAI score22

    Turing Award winners Diffie and Barto keynote BAAI Conference on AI security and RL

    AITuring Award winners Whitfield Diffie and Andrew Barto delivered keynotes at the BAAI Conference on AI security and reinforcement learning. Diffie argued that today's feedback-based approach only patches programs after they fail, and that formal methods offer a path to substantially more reliable intended behavior. Barto framed reinforcement learning around control, search, and associative memory, describing its core insight as caching search results rather than searching continuously.

Jun 12

Jun 12Fri

Jun 11

Jun 11Thu
  1. OpenRouter BlogAI score74

    OpenRouter Fusion panels beat individual models on the DRACO deep research benchmark

    AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.

    Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.

Jun 9

Jun 9Tue
  1. One Useful Thing (Ethan Mollick)AI score72

    Ethan Mollick tests Claude 5 Fable and finds it runs long projects with little user input

    AIEthan Mollick, who had early access to Claude 5 Fable, reports that it outperformed other public models in his tests, including an isochrone travel-time map and a nine-and-a-half-hour software build called Concord. He says the model delegated work to other agents and made many design choices he could not see or weigh in on, leaving him closer to a client than a hands-on operator. He also notes high token usage, frequent fallback to Claude 4.8 Opus under security guardrails, and persistent quirks in its writing style.

Jun 8

Jun 8Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score41

    Xiaomi releases MiMo-V2.5-Pro-FP4-DFlash, an FP4 model with block-diffusion decoding

    AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.

Jun 6

Jun 6Sat
  1. Ahead of AI (Sebastian Raschka)AI score32

    Raschka Lists 2026 LLM Research Papers from January Through May, Heavy on Reasoning and Efficiency

    AISebastian Raschka has published a curated list of LLM research papers he bookmarked from January through May 2026, not a complete survey of the field. The list is weighted toward reasoning models, reinforcement learning, and efficient inference, with added interest in agent harnesses, long context, and diffusion language models. He highlights Nvidia's Nemotron 3 Super, a 120B-A12B hybrid model alternating attention and Mamba-2 layers, as a must-read, and notes a 4B Nano variant and the 550B-A55B Nemotron 3 Ultra released two days earlier.

Jun 3

Jun 3Wed
  1. Mark ChenAI score25

    Mark Chen says OpenAI's models could match Mythos on cyber vulnerabilities

    AIOpenAI's Mark Chen said that after Mythos showed AI models can prove 80-year-old theorems, he expected them to also find cyber vulnerabilities, and they did. He added that researchers in math may now be thinking the same idea in reverse, applying cybersecurity-style capability to mathematics. The post offers no specific models, benchmarks, or figures.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

May 26

May 26Tue
  1. One Useful Thing (Ethan Mollick)AI score40

    Mollick Warns AI Writing Defaults Erode Learning and Human Thinking

    AIEthan Mollick argues that using AI as a default for writing, without thinking, risks undermining the human effort that builds skill and style. He cites two Wharton-linked studies: a Turkish high school experiment where ChatGPT access hurt test performance, and a Taipei Python course where a personalized AI tutor raised exam scores by 0.15 standard deviations. Mollick calls the difference how AI is used, not whether, and notes that the tools for tutor-style learning are not intuitive to access.

May 21

May 21Thu
  1. Mark ChenAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

May 19

May 19Tue
  1. koray kavukcuogluAI score72

    Google's Gemini 3.5 Flash beats Gemini 3.1 Pro on coding and agentic benchmarks

    AIGoogle's Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%). The post also claims it is 4x faster than other frontier models, or 12x in Antigravity, and reports 83.6% on MMMU-Pro for multimodal performance.

    Why it matters: The post gives specific benchmark scores against Gemini 3.1 Pro, letting readers compare coding, agentic, and multimodal results directly.

May 15

May 15Fri
  1. Intern Large ModelsAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

May 10

May 10Sun
  1. Thinking Machines LabAI score67

    Thinking Machines Lab previews interaction models for real-time human-AI collaboration

    AIThinking Machines Lab announced a research preview of interaction models that take in audio, video, and text continuously and respond in real time without external turn-detection harnesses. The model, TML-Interaction-Small, is a 276B-parameter MoE with 12B active parameters, paired with an asynchronous background model for sustained reasoning and tool use. The post reports competitive intelligence scores and lower turn-taking latency against GPT-realtime and Gemini Live models, along with new interactivity benchmarks where baseline models largely failed.

    Why it matters: The post explains a time-aligned, full-duplex design and benchmarks against turn-based models, showing how interaction and background reasoning can be split across two cooperating models.

May 9

May 9Sat
  1. PaddlePaddleAI score60

    Baidu releases ERNIE 5.1 with reduced pretraining cost and parameter scale

    AIBaidu's PaddlePaddle account announced ERNIE 5.1, which it says cuts total parameters to about one-third and activated parameters to about one-half, using roughly 6% of the pretraining cost of models at similar scale. The post reports benchmark results including 99.6 on AIME26 with tools, surpassing DeepSeek-V4-Pro on τ3-bench and SpreadsheetBench-Verified, and ranking #4 globally on Arena Search. ERNIE 5.1 is available through the ERNIE website and Baidu AI Studio Model Playground.

May 8

May 8Fri
  1. Berkeley AI ResearchAI score46

    Adaptive Parallel Reasoning Lets Models Decide When to Parallelize Inference

    AIBerkeley AI Research describes adaptive parallel reasoning, in which a reasoning model decides when to split independent subtasks, how many concurrent threads to spawn, and how to coordinate them. The approach targets the latency, context-rot, and cost problems of long sequential reasoning, which can require millions of tokens and tens of minutes for complex tasks. Existing methods such as self-consistency, Tree of Thoughts, ParaThinker, and Hogwild! Inference fix the parallel structure outside the model, which wastes compute on simple problems.

Apr 30

Apr 30Thu
  1. ARC PrizeAI score44

    GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models

    AIOpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs. The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic. ARC Prize is open-sourcing its analysis package.

Apr 27

Apr 27Mon
  1. Mistral AI · new models on Hugging FaceAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. ARC PrizeAI score58

    ARC Prize Releases Human Performance Dataset for ARC-AGI-3 Benchmark

    AIARC Prize Foundation released an open-source human dataset for ARC-AGI-3, covering 342 step-by-step replays across 25 public environments from a study of 458 participants. The source reports that every environment was solved by at least two humans, and it updates scoring by moving the per-level baseline to the median human player and raising the per-level cap from 100% to 115%.

Apr 9

Apr 9Thu
  1. Andrej KarpathyAI score45

    Karpathy says AI capability gap stems from uneven use and training

    AIAndrej Karpathy argues that people judging AI from free-tier ChatGPT or Advanced Voice Mode miss the strong capabilities of current agentic models like OpenAI Codex and Claude Code. He says gains are "peaky," concentrated in verifiable technical domains like programming and math that suit reinforcement learning and attract B2B investment, while writing and everyday advice improve less. Those who use frontier agentic tools professionally in these fields see far greater capability, which is why the two groups talk past each other.

Apr 1

Apr 1Wed
  1. Jim FanAI score62

    CaP-X open-sources agentic robotics toolkit, benchmark, and RL setup

    AIJim Fan announced the open-source release of CaP-X, an agentic robotics framework in which LLM-driven agents control robot arms and humanoids through perception and actuation APIs. The release includes CaP-Gym with 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR, and CaP-Bench, which evaluates 12 frontier LLMs and VLMs across 8 tiers. The post also reports that a 7B open-source model rose from 20% to 72% success after 50 RL training iterations, with synthesized programs transferring to real robots.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.