@googlecloud and Inferact are announcing today a partnership to make TPU a first-class citizen in @vllm_project.
This partnership puts both teams on one engineering roadmap to bring TPU to the broader open model ecosystem, optimizing vLLM as the agentic production serving engine for TPU:
- Production serving features and optimized kernels
- A native PyTorch path via TorchTPU
- Moving towards day-0 support for frontier model releases
We're also launching a community program: shared TPU capacity for open-source contributors, plus dedicated review and design help from the core vLLM maintainers at Inferact.
Everything this collaboration produces is open source. Read the full announcement: https://inferact.ai/news/google-tpu-partnership
