Skip to content
View original post on X: Zhihao Jia· 62/100AI score62/100

Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

AISummary

Lithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

Post on XView on X
@JiaZhihao

We’re open-sourcing lithos-metal 🚀

Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.

Ultra-fast inference on your laptop. Try it with any coding agent in one command.

Code: https://github.com/lithos-ai/lithos-metal
Tech blog: https://www.lithosai.com/blog/lithos-metal

Source: Zhihao Jia · x.comPublished · added here