Prime Inference launches serverless and reserved serving for open frontier models
Prime Inference: Fast, Reliable Serving for Frontier Open Models
AISummary
Prime Inference is a serving platform for frontier open-source models, offering serverless endpoints and reserved capacity on Prime's GPU infrastructure across multiple datacenters. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22, and the post reports a near-zero tool-call error rate and 100% uptime since launch. The post also describes GLM-5.3 serving on GB200 NVL72 with prefill/decode disaggregation and NVFP4 KV compression.
AIWhy it matters
The post separates scheduler, KV-cache, and tool-call fixes, showing concretely which bottlenecks shape production serving of open frontier models.
Source: Prime Intellect Blog · primeintellect.ai