Skip to content
Read the original: Prime Intellect· Published 38/100AI score38/100

Prime Intellect stores MLA KV cache in NVFP4 for more cached tokens

Original titleDecode: NVFP4 KV compression

AISummary

Prime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.

Read the original x.com

Source: Prime Intellect · x.comPublished · added here