Skip to content
Trending storyDeveloping

Lithos AI open-sources lithos-metal for local Qwen3.8-27B inference on Apple M5 Max

1 article1 sourcesince Oct 8Last article

Overview

AISummary of 1 article

Lithos AI says it has open-sourced lithos-metal, an inference tool that uses megakernels and DSpark speculative decoding to run Qwen3.8-27B locally.

According to the company's post, the setup reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. The company also says users can try it with any coding agent in one command, with code on GitHub and a technical blog linked. These performance figures are the company's own claims; the report does not include independent benchmarks.

Written by AI from the articles below · updated Oct 9, 8:28 AM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 8
  1. Zhihao Jia
    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post

Heat trend

Not enough continuous observations to show a trend yet.