Skip to contentSkip to stories

Updated

#Reasoning

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. SGLangOfficialAI score34

    SGLang reports inference speedups for MLA, MoE, and KDA kernels

    AISGLang says restructured MLA decode kernels on Rubin, which fit a deeper pipeline in 327 KiB of shared memory, deliver a 16% speedup at batch 16 with 128K context and bit-identical output. The post also reports 20% faster full FP8 MLA at batch 1 and 20% faster KDA verify kernels after keeping weights in registers and reducing synchronization. It additionally covers fusing MoE finalization, the shared expert, 8-GPU all-reduce, and RMSNorm into one collective kernel.

  2. SGLangOfficialAI score37

    Miles runs full RL loop on Rubin GPUs with SGLang and Megatron

    AIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.

  3. ARC PrizeOfficialAI score42

    ARC Prize 2026 ARC-AGI-2 high score reaches 88.06%

    AITufa Labs posted an 88.06% score on ARC-AGI-2, a new high for the ARC Prize 2026 leaderboard. ARC Prize says a $150K bonus prize, on top of guaranteed prizes, will be split among all teams scoring over 85%.

    Image from @arcprize's post
  4. QbitAINewsAI score67

    Aether AI shows CRIS-0 robot recovering from disturbances via causal reasoning

    AIAether AI, founded by UCSD assistant professor Biwei Huang, has released official demos of its CRIS-0 causal intelligence system for robots. In tests, the robot recovered from external disturbances in 9 of 10 random trials, typically within about 2 seconds, and stopped within 0.2 seconds when a human hand entered the workspace during a microwave-door task.

Oct 8

Oct 8Thu
  1. vLLMOfficialAI score62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    AIvLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

    Image from @vllm_project's post

Oct 6

Oct 6Tue
  1. Epoch AIOfficialAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    AIEpoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  2. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. ReflectionOfficialAI score23

    Reflection AI's Beam model pretrained in four weeks on 24T tokens

    AIReflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities. The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class. It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

    Image from @reflection_ai's post
  2. ReflectionOfficialAI score42

    Reflection AI previews Beam, a 500B open model under Apache 2.0

    AIReflection AI says its Beam model, with a 500B form factor, combines strong agentic performance and efficient reasoning for enterprises, governments, and developers. Beam is in final red-teaming and will be released this month under an Apache 2.0 license, with quantized FP8 and NVFP4 versions for efficient deployment. Early access sign-ups are open on the company's platform.

Oct 3

Oct 3Sat
  1. Aidan GomezXAI score46

    AlephAlpha releases Kolibri, a German-English model with a technical report

    AIAlephAlpha has released Kolibri, a German-English model with 78B total parameters, 3.46B active, and context up to 1M tokens. The weights are available under Apache 2.0 for running on users' own hardware. Cohere's Aidan Gomez congratulated the team on the model and its detailed technical report.

Oct 1

Oct 1Thu
  1. OpenRouter · New modelsBlogAI score36

    Apodex 1.1 Mini Released as Free Reasoning Model for Research Tasks

    AIApodex has released Apodex 1.1 Mini, a free reasoning-first model designed for complex, long-horizon research and forecasting tasks. According to the source, it works directly with files, data, code, and tools to produce verifiable results.

Sep 30

Sep 30Wed
  1. ollamaOfficialAI score34

    Ollama adds support for JEV-style decision models

    AIOllama now supports JEV-style decision models, with documentation for the decision capability and API. The post links to a blog announcement, decision capability docs, and API docs, but gives no further details on the models or their performance.

Sep 29

Sep 29Tue
  1. ModelScopeOfficialAI score54

    IQuest-Q1 released as 320B MoE model for long-horizon coding agents

    AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

    Image from @ModelScope2022's post