Skip to contentSkip to stories

Updated

#Reasoning

Showing low-relevance items too. Hide low-relevance items

Sep 2

Sep 2Wed
  1. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  2. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  3. Sebastian RaschkaAI score38

    Raschka Says OpenAI Astra's Looped Transformer Is Not a Big Deal

    AISebastian Raschka argues that the looped transformer approach attributed to OpenAI's Astra is a minor architectural tweak, not a major breakthrough. He explains that Nanbeige4.2-3B reuses its 22-layer stack twice, effectively doubling depth without adding weights but roughly doubling compute, and that the idea traces back to the Mixture-of-recursions NeurIPS paper. He adds that layer reuse does not inherently hide chain-of-thought, though it could shift more computation into latent activations.

    Image from @rasbt's post

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

Aug 30

Aug 30Sun
  1. Fireworks AI BlogAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

Aug 29

Aug 29Sat
  1. Tencent HyAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri

Aug 27

Aug 27Thu
  1. Leandro von WerraAI score22

    Pollen Robotics unveils Microduck, a $400 open-source RL biped robot

    AIPollen Robotics has unveiled Microduck, a 25 cm open-source biped with 15 actuators and sensors including a camera, speaker, and LiDAR that users can train with reinforcement learning. The robot ships with more than half a dozen pre-trained policies for walking, sitting, roller-skating, and picking up objects with its articulated beak, and costs less than $400.

Aug 26

Aug 26Wed

Aug 21

Aug 21Fri

Aug 19

Aug 19Wed

Aug 17

Aug 17Mon
  1. Jason WeiAI score45

    Jason Wei argues tool use cannot replace larger language models

    AIJason Wei now believes a small 1B-parameter "cognitive core" relying on tools is insufficient, because fast, natural recall without tool use matters. He cites speed, knowledge better learned through backpropagation than retrieved from search, and the greater reliability of already-known facts over repeated lookups. Since a 1B model has an information limit, he argues that demanding AI will still need larger models, not just tool access.

  2. Import AIAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 14

Aug 14Fri
  1. Epoch AI · The Epoch BriefAI score42

    Epoch AI lists nine big AI questions its benchmarks aim to answer

    AIEpoch AI outlines nine open questions about AI capabilities, including whether AI can take over full jobs and whether benchmark scores are correlated. The author says Epoch's benchmarking work is built to help answer them, citing examples such as MirrorCode, Remote Labor Index, and the Epoch Capabilities Index (ECI). The post notes that benchmark scores are highly correlated across domains, and that ECI growth trends can help detect whether AI capability progress has accelerated.

Aug 13

Aug 13Thu
  1. DeepSeekAI score62

    DeepSeek launches V4-Pro with Agent upgrades and OpenAI Responses API support

    AIDeepSeek announced the launch of DeepSeek-V4-Pro, citing major Agent upgrades and flexible reasoning effort settings of low, high, and max for V4-Pro and V4-Flash. The model supports the native OpenAI Responses API and is optimized for Codex with one-click setup. V4-Pro is available on the app and web through Expert Mode and via API, with model names unchanged.

    Image from @deepseek_ai's post
  2. Sebastien BubeckAI score51

    Neurosurgery resident uses ChatGPT 5.6 to prove Crouzeix's conjecture

    AIA neurosurgery resident at Peking Union Medical College Hospital, Shanmu Jin, posted a preprint claiming a proof of Crouzeix's conjecture in numerical linear algebra after using ChatGPT 5.6. The essay by Alex Townsend and Anne Greenbaum says the conjecture had been open for more than two decades and that the argument held up after a few hours of review. Bubeck, who says he spent a week on the problem in 2012, quotes the story and calls it amazing.

Aug 12

Aug 12Wed
  1. Tri DaoAI score36

    Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

    AITri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

  2. Michael TruellAI score62

    Grok 4.6 is released with gains on agentic and knowledge-work benchmarks

    AIGrok 4.6 is released as a significant improvement over Grok 4.5 at the same price, according to the announcement. The author says it is significantly better at difficult tasks and knowledge work, combining Opus-class intelligence and polish with low cost and high speed. A comparison table shows Grok 4.6 High scoring 61 on the AA Intelligence Index, versus 56 for Grok 4.5 High, and 1753 on GDPval-AA v2, versus 1526.

Aug 11

Aug 11Tue
  1. Fireworks AI BlogAI score45

    Fireworks AI Tests Anthropic's J-Lens on Kimi K3 and Qwen3.5-9B

    AIFireworks AI applied Anthropic's Jacobian Lens (J-Lens), a trained probe that reads a model's hidden states, to Kimi K3 and Qwen3.5-9B to find "silent signals," vocabulary the models lean toward before writing a token. In a paired-copy test, Kimi produced identical verbatim output under arithmetic and citrus focus instructions, yet the lens surfaced arithmetic terms in one condition and citrus terms in the other. Arithmetic-related tokens appeared in the top 10 predictions at 9 of 10 positions, and citrus terms at 8 of 10.

Aug 10

Aug 10Mon
  1. Chip HuyenAI score28

    Chip Huyen jokes about sending instructions in all caps

    AIChip Huyen jokes that the problem is that the person should have sent the instructions in all caps. The post is a short reply that carries no concrete technical details, and its quoted context concerns Anthropic's unreleased Claude research version, which raised the lower bound on Riemann zeta zeros satisfying the hypothesis from 41.6% to 67.2%.

    Image from @chipro's post

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. Prime Intellect BlogAI score62

    Prime Intellect adds multi-agent training and evaluation to PRIME-RL

    AIPrime Intellect's RL stack now supports multi-agent systems, letting users program interactions between agents, choose which roles learn, and assign credit across an episode. The release introduces Agent and Env abstractions and four example patterns: agentic judging, self-play, and user simulation. Multi-agent support ships today in verifiers 0.3.0 and prime-rl 0.8.0.

    Why it matters: The post explains the Agent and Env abstractions and four multi-agent patterns, showing how roles, credit assignment, and episodes can be programmed in one RL stack.

Aug 6

Aug 6Thu

Aug 3

Aug 3Mon

Aug 1

Aug 1Sat
  1. Sebastien BubeckAI score78

    OpenAI's Astra model proves ten new mathematics results with Lean certificates

    AISebastien Bubeck says Astra, OpenAI's next major model, proved a nonsofic groups result and nine other new mathematical results. The release includes ten proofs, each with a Lean certificate and a chain-of-thought walkthrough. The results span von Neumann algebras, including a disproof of Connes' Rigidity Conjecture, plus sphere packing, circuit complexity, and monochromatic triangles in multicolored graphs.

    Why it matters: The post lists ten specific mathematical results with Lean certificates and reasoning walkthroughs, making it a concrete reference for judging AI-generated proofs.

Jul 30

Jul 30Thu