Prime Intellect stores MLA KV cache in NVFP4 for more cached tokens
Original titleDecode: NVFP4 KV compression
AISummary
Prime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.
Source: Prime Intellect · x.comPublished · added here