Lithos AI open-sources lithos-metal for local Qwen3.8-27B inference on Apple M5 Max
Overview
Lithos AI says it has open-sourced lithos-metal, an inference tool that uses megakernels and DSpark speculative decoding to run Qwen3.8-27B locally.
According to the company's post, the setup reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. The company also says users can try it with any coding agent in one command, with code on GitHub and a technical blog linked. These performance figures are the company's own claims; the report does not include independent benchmarks.
Written by AI from the articles below · updated Oct 9, 8:28 AM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Zhihao Jia@JiaZhihaoLithos AI open-sources lithos-metal for fast local inference on Apple M5 MaxAILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

Heat trend
Not enough continuous observations to show a trend yet.