Skip to content
Read the original: IThome · AI· Published 41/100AI score41/100

Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s

Original title (Chinese)

12GB 显存显卡跑 125B Qwen3.8 模型:Strata 登场,单张 RTX 5070 跑出 94 词元 / 秒

AISummary

Developer Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding.

On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.

Read the original ithome.com

Source: IThome · AI · ithome.comPublished · added here