News · September 8, 2026
Inferact Optimizes vLLM for Real-World Agentic Serving, Achieving Top Efficiency Measured on Blackwell GPUs
Agentic workloads are becoming a major source of vLLM traffic. Their long-running, multi-turn sessions combine long contexts, short outputs, and extensive prefix reuse, which demands optimizations across the entire serving stack. Inferact led a coordinated effort with the vLLM community spanning KV cache management, parallelism and engine optimizations, and prefill/decode disaggregation.
Measured on AgentX, SemiAnalysis’s public agentic benchmark, vLLM achieves up to 130K total tokens per GPU-second on DeepSeek V4 Pro running on NVIDIA GB300, and an interactivity of up to 376 tokens per second on MiniMax M3 running on NVIDIA B300.
Read the full technical deep dive on the vLLM blog: vLLM x AgentX: Optimizing for Real-World Agentic Serving.