Skip to content

Areas · Latest news

China AI

What Chinese labs and companies ship and how China regulates AI: DeepSeek, Qwen, Kimi, GLM, MiniMax and others, plus chips and policy.

69 top picks · 18 in the past 30 days · chosen from 494 items collected

Latest pick

Top picks archive · Page 3

Top picks 41–60 of 69

Jul 28

Jul 28Tue
  1. MiniMax · new models on Hugging FaceAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

Jul 27

Jul 27Mon
  1. KimiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

Jul 23

Jul 23Thu
  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

Jun 16

Jun 16Tue
  1. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. Z.ai Release NotesAI score62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    AIZ.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 11

Jun 11Thu
  1. OpenRouter BlogAI score74

    OpenRouter Fusion panels beat individual models on the DRACO deep research benchmark

    AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.

    Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.

  2. Moonshot AI (Kimi) · new models on Hugging FaceAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 26

May 26Tue
  1. MiniMax BlogAI score67

    MiniMax Agent Team Adds Parallel Multi-Agent Collaboration for Long Tasks

    AIMiniMax has upgraded its Agent, renamed Mavis, and introduced Agent Teams that run multiple role-based Agents in parallel on desktop. The team uses Leader, Worker, and Verifier roles so complex tasks can be split, checked, and reported at key checkpoints, and it merges TokenPlan and Agent Plan into one subscription with credits shared between Agent and API. The post also discusses the added token, handoff, and retry costs of multi-Agent work, and says the Agent will be open-sourced alongside MiniMax M3.

    Why it matters: The post explains why multi-Agent helps long tasks and where its verification, token, and aggregation costs come from, useful for judging when a team setup beats a single Agent.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens

    AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.

    Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

  2. DeepSeek API NewsAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    AIThe DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    Why it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Mar 17

Mar 17Tue
  1. MiniMax BlogAI score63

    MiniMax M2.7 takes part in its own model and harness evolution

    AIMiniMax says M2.7 is its first model to deeply participate in its own evolution, building agent harnesses and running reinforcement learning experiment workflows. The post reports 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, and a 30% improvement on an internal evaluation set after more than 100 autonomous optimization rounds. It also states that M2.7 handles 30%-50% of its research team's workflow, though human researchers still make critical decisions.

    Why it matters: The post ties M2.7's self-evolution claims to specific benchmark numbers and workflow details, helping readers judge how much of the iteration loop is autonomous.

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.