Skip to contentSkip to stories

Updated

#Reasoning

Items with an AI score under 20 are hidden. Show low-relevance items

Mar 4

Mar 4Wed
  1. Tri DaoAI score62

    Tri Dao Shares Speculative Speculative Decoding, a Claimed Up-to-2x LLM Inference Speedup

    AITri Dao reposts a quoted post from @tanishqkumar07 introducing Speculative Speculative Decoding (SSD), an LLM inference algorithm claimed to be up to 2x faster than leading inference engines. The quoted post credits collaborators @tri_dao and @avnermay and links to a thread with details. Tri Dao's own text says the approach applies an asynchronous-machines principle seen in GPU kernels to speculative decoding.

  2. Mistral AI · new models on Hugging FaceAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 combines instruct, reasoning, and Devstral capabilities in one multimodal model with 119B total parameters, 6.5B active per token, and a 256k context window. The source reports a 40% reduction in latency-optimized end-to-end completion time and 3x more requests per second in throughput-optimized setups versus Mistral Small 3. It is released under Apache 2.0 and supports reasoning mode toggling per request.

    Why it matters: The source lists architecture, context length, and mode-switching controls, letting readers compare this release's design with earlier Mistral Small models.

Feb 28

Feb 28Sat
  1. Cognition Blog (Devin, Windsurf)AI score36

    Cognition Previews SWE-1.6, Claims 11% Gain Over SWE-1.5 on SWE-Bench Pro

    AICognition previewed its ongoing SWE-1.6 training run, which scores 11% higher than SWE-1.5 on SWE-Bench Pro and runs at 950 tok/s. The model is post-trained on the same pre-trained model as SWE-1.5, and the company is rolling out early access to a small group of users to gather feedback on behavior such as overthinking and excessive self-verification. The company says training steps now run 6x faster than three months ago, with rollouts in NVFP4 precision.

Feb 25

Feb 25Wed
  1. Quoc LeAI score53

    Google's Aletheia math agent solves 6 of 10 FirstProof problems

    AIQuoc Le announced that Aletheia, a math research agent, autonomously solved 6 of 10 FirstProof problems, the best result in the inaugural challenge. The post says this exceeds last year's IMO-gold achievement and points to a paper and thread for full details. The accompanying figure shows 10 unmodified problems, 6 candidate solutions per agent, and expert evaluation yielding 6 solved problems on a best-of-2 basis.

  2. Quoc LeAI score65

    Aletheia Agent Solves 6 of 10 FirstProof Math Problems Autonomously

    AIGoogle researchers used the Aletheia agent, powered by Gemini 3 Deep Think, to attempt 10 FirstProof challenge problems without modification. The agent operated fully autonomously and solved 6 of the 10 problems, according to the post, with methodology and expert evaluations described in the linked arXiv paper.

    Why it matters: The post gives the autonomous setup and expert-evaluated results for an AI agent on FirstProof math problems, useful for judging how far such systems go on research-level math.

    Image from @quocleix's post

Feb 19

Feb 19Thu
  1. Yi TayAI score78

    Google releases Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2

    AIGoogle has released Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2 and more than twice the score of Gemini 3 Pro on that benchmark. The model is rolling out to developers in preview through the Gemini API and Google AI Studio, to enterprises via Vertex AI and Gemini Enterprise, and to consumers in the Gemini app and NotebookLM.

    Why it matters: The post pairs the release with a benchmark table comparing Gemini 3.1 Pro against Gemini 3 Pro, Claude Sonnet 4.6, Claude Opus 4.6, and GPT-5.2 on reasoning and coding tasks.

Feb 14

Feb 14Sat

Feb 13

Feb 13Fri
  1. Jakub PachockiAI score62

    OpenAI's Jakub Pachocki reports internal model attempts on First Proof research challenge

    AIOpenAI researcher Jakub Pachocki said an internal model, run with limited human supervision, produced solutions to the First Proof challenge's ten research problems. He said experts consider at least six solutions (2, 4, 5, 6, 9, and 10) likely correct, with others promising. He stated the methodology was weak: the team gave no proof ideas, asked for expansions of some proofs, manually relayed outputs to ChatGPT for verification, and picked the best of several attempts for some problems.

  2. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 12

Feb 12Thu

Feb 11

Feb 11Wed
  1. Yi TayAI score67

    Aletheia math research agent produces two papers and solves open Erdős problems

    AIYi Tay introduces Aletheia, a math research agent powered by an advanced version of Gemini Deep Think. The post says it produced two publishable papers, one fully automatic and one human-AI collaboration, and solved multiple open Erdős problems. The attached image shows a Google DeepMind paper titled "Towards Autonomous Mathematics Research" with a generator, verifier, and reviser loop.

    Image from @YiTayML's post

Feb 10

Feb 10Tue
  1. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5, a 744B-parameter open model for agentic engineering

    AIZ.ai launches GLM-5, scaling from 355B to 744B total parameters with 40B active and pre-training data from 23T to 28.5T tokens. The model integrates DeepSeek Sparse Attention to reduce deployment cost and reports strong results on reasoning, coding, and agentic benchmarks against GLM-4.7, DeepSeek-V3.2, Kimi K2.5, and several frontier models.

    Why it matters: The source gives concrete scale, data, and benchmark comparisons against named frontier models, showing where GLM-5 sits among open-source and proprietary systems.

Feb 2

Feb 2Mon

Jan 23

Jan 23Fri
  1. Mistral AI · new models on Hugging FaceAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 is a 119B-parameter MoE model with 6.5B active per token and a 256k context window, combining instruct, reasoning, and Devstral-style coding in one model. It accepts text and image input, lets users set reasoning_effort per request, and is released under Apache 2.0. The model card reports a 40% latency reduction and 3x throughput versus Mistral Small 3 in its tested setups, and its benchmark chart shows reasoning scores on GPQA Diamond, MMLU Pro, AIME-style text tasks, and MMMU-Pro.

    Why it matters: The model card names concrete architecture, context, and licensing details, letting readers compare its reasoning toggle and efficiency claims against other open models.

Jan 19

Jan 19Mon
  1. Z.ai (GLM) · new models on Hugging FaceAI score62

    Z.ai releases GLM-4.7-Flash, a 30B-A3B MoE model for lightweight deployment

    AIZ.ai has released GLM-4.7-Flash, a 30B-A3B MoE model that it positions as the strongest model in the 30B class. The model reports SWE-bench Verified 59.2 and τ²-Bench 79.5, and supports local deployment through vLLM and SGLang.

    Why it matters: The source lists benchmark scores against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B, letting readers compare the 30B-class MoE model directly with its named rivals.

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

Dec 17, 2025

Dec 17, 2025Wed

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

    Image from @nickaturley's post

Dec 10, 2025

Dec 10, 2025Wed
  1. Tim DettmersAI score60

    Tim Dettmers argues AGI will not happen due to physical computing limits

    AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

  2. Yi TayAI score38

    Google DeepMind's Gemini team launches new reasoning research group in Singapore

    AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.

    Image from @YiTayML's post

Nov 29, 2025

Nov 29, 2025Sat
  1. Andrej KarpathyAI score62

    Karpathy argues LLMs are a new kind of intelligence shaped by commercial, not evolutionary, pressure

    AIKarpathy argues animal intelligence is only one point in a large space of possible minds, and LLMs arise from a fundamentally different optimization process. He contrasts survival-driven animal drives with LLM training shaped by imitation of human text, RL on task distributions, and user engagement metrics, which he says leaves LLMs jagged and prone to sycophancy. He calls LLMs humanity's first contact with non-animal intelligence and says people who build accurate internal models of them will reason about them better.

Nov 28, 2025

Nov 28, 2025Fri

Nov 22, 2025

Nov 22, 2025Sat

Nov 18, 2025

Nov 18, 2025Tue

Nov 17, 2025

Nov 17, 2025Mon
  1. Andrej KarpathyAI score60

    Karpathy argues verifiability predicts which tasks AI automates fastest

    AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Oct 27, 2025

Oct 27, 2025Mon
  1. Lilian WengAI score44

    On-policy distillation uses a teacher model as dense process reward

    AILilian Weng says on-policy distillation lets a teacher model act as a process reward model, providing dense rewards during training. The approach also prevents the out-of-distribution shock that SFT-style training can cause during rollouts. Thinking Machines' related post reports it outperforms other approaches for math reasoning and an internal chat assistant at a fraction of the cost.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.