Skip to content
View original post on X: Inferact· 42/100AI score42/100

Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

AISummary

Inferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

Post on XView on X
@inferact

Behind this blog is months of our team's work tuning vLLM on agentic workloads and validating on @SemiAnalysis_ AgentX benchmark. We find that open source models optimized for agentic workloads reach up to 130K tokens/GPU-sec, 106× cheaper than Opus 5 API pricing.

vLLM is the open source agentic production serving engine. Inferact optimizes vLLM and builds enterprise inference on top of it. 🚀

vLLM@vllm_project
New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime optimizations, measured on AgentX, @SemiAnalysis_'s public agentic benchmark. 🧵1/6
View quoted post on X

Source: Inferact · x.comPublished · added here