Skip to contentSkip to stories

Updated

#Reasoning

Showing low-relevance items too. Hide low-relevance items

Sep 30

Sep 30Wed
  1. Artificial Analysis ArticlesAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    AIKorean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  2. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    AIArtificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    Why it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

Sep 29

Sep 29Tue
  1. Fireworks AI BlogAI score51

    Fireworks explains how numerical mismatch and MoE routing can derail RL training

    AINumerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

  2. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  3. Jerry LiuAI score22

    Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments

    AIJerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks. The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult. The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.

    Image from @jerryjliu0's post
  4. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

    Image from @OpenBMB's post
  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score40

    InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs

    AIInternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.

Sep 28

Sep 28Mon
  1. Google · Gemini appAI score38

    See what 4 builders are making with Gemini 3.8 Flash

    AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.

Sep 27

Sep 27Sun
  1. Sakana AIAI score46

    Sakana AI's SAIL boosts VLM robot trajectory success via test-time scaling

    AISakana AI and the University of Tokyo introduced SAIL, a method that generates robot trajectories with a VLM and refines them through simulator testing, VLM feedback, and Monte Carlo tree search. Across six simulated manipulation tasks, raising the search budget from one candidate to 45 increased the success rate of finding a working trajectory from 25% to 73%. The authors also tested the approach on a physical robot, though the post frames further transfer to real hardware as an open question.

    Video from @SakanaAILabs's post
  2. Alexander DoriaAI score18

    Doria Argues Local Language Models Ease Industrial Audit Demands

    AIAlexander Doria argues that embedding AI models in industrial processes requires easing audits and supporting local languages and professional registers. He says this need is real despite sarcasm about the idea. Background from Aleph Alpha notes that reasoning models think in English even for German prompts, and that small German reasoning data doses can hurt performance while large doses mostly recover it.

  3. Tibor BlahoAI score85

    OpenAI releases GPT-6 Sol and Luna as Anthropic launches Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.

    Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.

    Video from @btibor91's post
  4. Xiaomi MiMoAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. Sebastian RaschkaAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

    Video from @rasbt's post
  2. Alexander DoriaAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

Sep 25

Sep 25Fri
  1. Kevin Weil 🇺🇸AI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

  2. AnthropicAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 24

Sep 24Thu
  1. Goodfire ResearchAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    AIGoodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

  2. Goodfire ResearchAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  3. Goodfire ResearchAI score57

    Goodfire finds sparse autoencoder features capture curved neural geometry in three ways

    AIGoodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them. The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution. It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.

Sep 23

Sep 23Wed
  1. Tencent HyAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Redwood Research BlogAI score71

    Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues

    AIRedwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.

  3. Black Forest LabsAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

    Image from @bfl_ai's post

Sep 22

Sep 22Tue
  1. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  3. Sebastian RaschkaAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  4. Interconnects (Nathan Lambert)AI score34

    Epoch AI's JS Denain Debates RSI, US-China Gap, and AI Jaggedness

    AIJS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.

Sep 21

Sep 21Mon
  1. Latent.SpaceAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  2. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  3. Xiaomi MiMoAI score44

    MiMo-V2.6-Pro assists scientific research in materials and formal mathematics

    AIXiaomi's MiMo-V2.6-Pro, without research-specific RL training, helped Xiaomi materials researchers propose MOF materials for capturing PFAS "forever chemicals" and ran computational screening for wet-lab validation. It also helped formalize the full main theorem of Li–Yorke's "Period Three Implies Chaos" in Lean 4, producing a project of 6,000+ lines verified by Lean's kernel with no unfinished proof placeholders.

    Video from @XiaomiMiMo's post