These changes land in open-source vLLM, and the Inferact platform benefits directly from that upstream tuning. Founded by the creators and core maintainers of vLLM, Inferact runs open models in production on your compute or ours.
Early work on hardware like NVIDIA Vera Rubin shows us where inference is heading, and we use what we learn to optimize inference for your workload.
Learn more at http://inferact.ai
