Skip to content
View original post on X: Inferact· 49/100AI score49/100

Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

AISummary

Inferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline.

The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding.

Inferact says it is open-sourcing the kernel today.

Post on XView on X
@inferact

Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding.

To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8.

We are open sourcing it today.

1/2

Source: Inferact · x.comPublished · added here