Skip to content
Read the original: IThome · AI· 46/100AI score46/100

JetBrains Releases Mellum2.1 Coding Model With Near-Double Qwen3.5-9B Throughput

Original title (Chinese)

JetBrains 编程 AI 模型 Mellum2.1 发布:高负载推理吞吐量近 Qwen3.5-9B 两倍

AISummary

JetBrains released Mellum2.1, a 12B mixture-of-experts coding model with 2.5B active parameters under Apache 2.0, emphasizing agentic programming.

Under high load, its inference throughput in tokens is nearly twice that of Qwen3.5-9B in JetBrains' comparison, and multi-token prediction (MTP) speeds single-request responses by about 1.6x.

The model is available on Hugging Face for local or private-infrastructure deployment, with GGUF and vLLM MTP support announced for later.

Read the original ithome.com

Source: IThome · AI · ithome.comPublished · added here