Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM
Qwen3.8-Flash can now be run locally! 🔥
AISummary
Unsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).
AIWhy it matters
The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.
Source: Unsloth AI · x.com