Skip to contentSkip to stories

Updated

#Reasoning

Showing low-relevance items too. Hide low-relevance items

Oct 2

Oct 2Fri
  1. AI at MetaAI score22

    Muse Spark helps prove finite-time blow-up in a laser-inspired wave model

    AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

    Image from @AIatMeta's post
  2. AI at MetaAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

  3. Redwood Research BlogAI score34

    Capabilities research pushes the safety-usefulness frontier too, not just safety research

    AIThe post argues that counting all research as safety work because it widens the safety-usefulness Pareto frontier is misleading. Safety research typically creates new safety options without boosting usefulness, while capabilities research typically raises usefulness at safety's expense, so developers tend to choose less safe points.

  4. Harrison ChaseAI score53

    Google Research's Cogentic uses multi-agent proof search to produce verified results

    AIGoogle Research's Cogentic is a multi-agent harness running on Gemini that searches for proofs of open theoretical computer science problems without expert hints. It runs rounds where an orchestrator launches provers, two adversarial verifiers must both accept each draft, and shared disk documents store attempts and verified lemmas. The system produced new results on five open problems in online learning, auction theory, and mechanism design, each checked by domain experts.

  5. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  6. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

    Image from @huggingface's post
  7. MIT Technology Review · AIAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    AIThore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.

Oct 1

Oct 1Thu
  1. François CholletAI score62

    Chollet Argues Reasoning Models Differ from Base LLMs by Inductive Program Prediction

    AIFrançois Chollet argues the key difference between base LLMs and modern LRMs is a shift from transductive answer prediction to inductive prediction of the program or reasoning chain behind an answer. He says this enables test-time induction and substantial fluid intelligence in LRMs, which he claims base LLMs largely lack. He cites ARC 1 results: base LLMs remain around 10-15%, while LRMs of the same size or smaller saturated the benchmark in 2025.

  2. Alexander DoriaAI score54

    SYNTH paper proposes fully synthetic single-stage training for reasoning models

    AIThe SYNTH paper, titled It's All Training, presents a fully synthetic single-stage pipeline for training workable reasoning models with high data efficiency. The authors argue this approach does not separate training into pretraining, mid-training, or post-training stages. The image shows the paper's abstract, which describes a pipeline built from a 58,000-article Wikipedia-based synthetic corpus and models named Baguettotron-600M and Baguettotron-MoE.

    Image from @Dorialexander's post
  3. Amazon ScienceAI score34

    Amazon Science Explains Graph-Centric Agentic AI for Network Root Cause Analysis

    AIAmazon Science describes a graph-centric approach in which a network digital twin graph and cascaded graph algorithms, orchestrated by an agentic AI layer, identify root causes in complex network failures. The approach was demonstrated with NTT DOCOMO at the Mobile World Conference, achieving root cause analysis in minutes on commercial networks. The article traces how graphs evolved from topology models to active reasoning substrates for agents.

  4. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

Sep 30

Sep 30Wed
  1. Apple Machine Learning ResearchAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  2. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  3. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  4. Tencent HyAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    AIResearchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

    Image from @TencentHunyuan's post
  5. OpenBMBAI score42

    Diffusion Reward Models learn full human preference distributions, not single scores

    AIOpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score. The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs. DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.

    Image from @OpenBMB's post