Skip to contentSkip to stories

Updated

#Reasoning

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 26

Sep 26Sat
  1. Sebastian RaschkaAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

    Video from @rasbt's post
  2. Alexander DoriaAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

Sep 25

Sep 25Fri
  1. Kevin Weil 🇺🇸AI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

  2. AnthropicAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 24

Sep 24Thu
  1. Goodfire ResearchAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    AIGoodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

  2. Goodfire ResearchAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  3. Goodfire ResearchAI score57

    Goodfire finds sparse autoencoder features capture curved neural geometry in three ways

    AIGoodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them. The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution. It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.

Sep 23

Sep 23Wed
  1. Tencent HyAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Redwood Research BlogAI score71

    Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues

    AIRedwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.

  3. Black Forest LabsAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

    Image from @bfl_ai's post

Sep 22

Sep 22Tue
  1. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  3. Sebastian RaschkaAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  4. Interconnects (Nathan Lambert)AI score34

    Epoch AI's JS Denain Debates RSI, US-China Gap, and AI Jaggedness

    AIJS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.

Sep 21

Sep 21Mon
  1. Latent.SpaceAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  2. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  3. Xiaomi MiMoAI score44

    MiMo-V2.6-Pro assists scientific research in materials and formal mathematics

    AIXiaomi's MiMo-V2.6-Pro, without research-specific RL training, helped Xiaomi materials researchers propose MOF materials for capturing PFAS "forever chemicals" and ran computational screening for wet-lab validation. It also helped formalize the full main theorem of Li–Yorke's "Period Three Implies Chaos" in Lean 4, producing a project of 6,000+ lines verified by Lean's kernel with no unfinished proof placeholders.

    Video from @XiaomiMiMo's post
  4. Jeff DeanAI score30

    Jeff Dean thanks Dawn Song after discussing AI's future

    AIJeff Dean, who recently left Google after 27 years, thanked Dawn Song for a discussion covering foundational ideas, recursive self-improvement, automated scientific discovery, and AI safety. The post is a brief acknowledgment of that conversation, which Song promoted as Dean's first public talk since leaving Google.

  5. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  6. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  7. howie.seriousAI score34

    Agrees with critique that GPT-6 Astra lags on open-ended tasks

    AIResponding to a post by ScarletKc, howie.serious simply agrees with the claim that GPT-6 Astra struggles with open-ended, exploratory work that lacks a fixed correct answer. The main post is a one-word endorsement (), while the quoted post argues GPT models excel at verifiable, goal-defined tasks and that Claude Fable handles open-ended exploration better.

Sep 19

Sep 19Sat
  1. Interconnects (Nathan Lambert)AI score47

    Why Nathan Lambert Still Doubts True Recursive Self-Improvement in AI

    AINathan Lambert argues that frontier labs such as OpenAI and Anthropic, which run thousands of concurrent agents, are amplifying anxiety about AI risk and progress. He says automatable research is too narrow to produce a large net acceleration, citing exponential scaling-law costs, diminishing returns from parallel agents, and resource bottlenecks. He would revise his view only if labs achieved unpredictable foundational breakthroughs.

  2. Sebastian RaschkaAI score36

    Raschka's Inference Scaling Part 1: Sampling for Better Accuracy

    AISebastian Raschka starts a series on inference scaling by modifying text generation with temperature scaling, top-p filtering, and multinomial sampling to produce diverse outputs. He says this enables self-consistency and best-of-N approaches that improve answer accuracy by more than 2x. The video covers chain-of-thought prompting, a MATH-500 evaluation, and accuracy versus compute tradeoffs.

    Video from @rasbt's post

Sep 18

Sep 18Fri
  1. TinkerAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

Sep 17

Sep 17Thu