Skip to content
Read the original: Awni Hannun· Published 51/100AI score51/100

Mirai releases speculative decoding in Uzu for Qwen3.6-27B on Apple M5 Max

Original titleQwen 27B dense at 105 tok/s output on an m5 max is pretty bonkers. Breaking down the memory wall one brick at a time.

AISummary

Mirai is releasing speculative decoding in its Uzu inference engine, starting with Qwen3.6-27B. The quoted post reports 105 output tokens per second on an Apple M5 Max with 128 GB of unified memory, 2.9× faster than the fastest MLX speculative-decoding implementation Mirai benchmarked.

The stack combines DFlash with Mirai's Weaver model, tree-based speculative decoding, Mirai quantization, and Metal kernels for Apple silicon.

Read the original x.com

Source: Awni Hannun · x.comPublished · added here