Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀
Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels.
Read the technical deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization
