LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max
Original titleQwen3.8-27B, up to 144 tok/sec on M5 Max.
AISummary
LM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.
Source: LM Studio · x.comPublished · added here