Skip to contentSkip to stories

Updated

#Open-source ecosystem

Sep 21

Sep 21Mon
  1. Tencent HunyuanAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  2. vLLM BlogAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  3. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

  4. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  5. Xiaomi MiMo · new models on Hugging FaceAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  6. LMSYS OrgAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

  7. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  8. ModelScopeAI score3

    ModelScope Invites Visitors to Apsara 2026 Booth for Merch

    AIApsara 2026 is ON 🔥 Come find the ModelScope booth, snap a pic, and take home some merch: 🎒 Backpacks · Crossbody bags · Tote bags 🛋️ Neck pillows · ☕ Mugs · 💻 Laptop stands 🧸 Plush pendants · 🧢 Hats · 🧲 Fridge magnets 🎧 Earphone pouches · 📱 Phone chains ...and more. 📍 Hangzhou International Expo Center Phase II · 1F · Intelligence Engine · Booth 1-5C · ModelScope Come early before they're gone 👀

  9. Tim DettmersAI score62

    Tim Dettmers argues academic labs can lead research through open local AI tools

    AITim Dettmers argues that academic labs can do their most important AI research by building coherent open-source ecosystems rather than competing on GPU scale. He describes his lab's upcoming open-source week, including an agent harness that optimizes kernels autonomously, local inference of large Qwen and DeepSeek models on consumer hardware, and an auto-compaction technique called CliffCompaction that he says cuts costs by about fifty percent.

  10. Interconnects (Nathan Lambert)AI score65

    Chinese labs lead open-weight models in benchmarks, downloads, and research use

    AINathan Lambert argues that Chinese open-weight models now lead American ones on benchmarks, Hugging Face downloads, and OpenRouter usage. He estimates the gap to the American closed frontier at 2 to 5 months for Chinese open models and 6 to 9 months for American open models. The piece also reports that Chinese open-weight models were mentioned in over 40% of arXiv papers he scanned, compared with 30% for American models.

Sep 20

Sep 20Sun
  1. ModelScopeAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

  2. Qwen · new models on Hugging FaceAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    Why it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

Sep 19

Sep 19Sat
  1. StepFunAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

Sep 18

Sep 18Fri
  1. LM StudioAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

  2. LMSYS OrgAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

  3. SemiAnalysisAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

  4. The Register · AIAI score34

    KDE turns 30 as Akademy weighs an AI-native desktop proposal

    AIKDE's Akademy conference in Graz, Austria, opens on September 19, where contributors Eva Brucherseifer and Jan Muehlig will present a talk proposing an "AI-native" KDE desktop built on a personal, encrypted "Kadai" kernel. The proposal's middle section is expected to divide attendees, while the project marks its 30th anniversary, with KDE 1.0 released in July 1998.

  5. Liquid AI · new models on Hugging FaceAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

Sep 17

Sep 17Thu
  1. vLLM BlogAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  2. Google · AI blogAI score38

    UN System Data Commons unifies global statistics into an AI-ready open platform

    AIThe United Nations system launched UN System Data Commons, an open-source platform built on Data Commons by Google that integrates siloed global statistics into one AI-ready knowledge graph. Users can query it in natural language, browse by location or theme, and use MCP-enabled AI agents to fetch verified figures and draft charts or reports. The UN plans to add more datasets, aiming to include 80% of UN system statistical datasets by 2027.

  3. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

  4. OpenBMBAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

  5. inclusionAI (Ant Ling) · new models on Hugging FaceAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  6. Ai2 (Allen Institute for AI)AI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

  7. inclusionAI (Ant Ling) · new models on Hugging FaceAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.