Mirai releases speculative decoding in Uzu for Qwen3.6-27B on Apple M5 Max
AIMirai is releasing speculative decoding in its Uzu inference engine, starting with Qwen3.6-27B. The quoted post reports 105 output tokens per second on an Apple M5 Max with 128 GB of unified memory, 2.9× faster than the fastest MLX speculative-decoding implementation Mirai benchmarked. The stack combines DFlash with Mirai's Weaver model, tree-based speculative decoding, Mirai quantization, and Metal kernels for Apple silicon.



