Modal trains DFlash speculator, faster than MTP for inference
Original titleModal trained a DFlash speculator that's much faster than MTP, making it a great boost for inference speeds!
AISummary
Modal has trained a DFlash speculator that runs much faster than MTP, according to Soumith Chintala. The speculator is backed by Inkling by Thinking Machines, which Modal says delivers 67% higher throughput and interactivity on Modal Auto Endpoints with SGLang.
Source: Soumith Chintala · x.comPublished · added here