Skip to content
Read the original: LM Studio· Published 62/100AI score62/100

LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

Original titleQwen3.8-27B, up to 144 tok/sec on M5 Max.

AISummary

LM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

Read the original x.com

Source: LM Studio · x.comPublished · added here