←Back to NewsAI News/InfranewsInfraGpusSGLang Squeezes 78% More AI Throughput on NVIDIA's Blackwell GPUsSGLang, Qwen, and NVIDIA ship 4-bit NVFP4 KV cache on Blackwell, packing 1.78x more context and boosting long-context decode up to 78%.SourceLMSYS OrgPublishedSep 21, 2026, 4:44 PMAuthorAlphaSignal NewsroomRead1 min readSGLang, Qwen, and NVIDIA ship 4-bit NVFP4 KV cache on Blackwell, packing 1.78x more context and boosting long-context decode up to 78%.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsAlphaSignal · repoInco AI's Splash Runs Qwen3.8-27B Twice as Fast on Apple SiliconTogether AI · newsTogether AI's ThunderKittens Hits 22.4 PFLOPS on NVIDIA's RubinIntelligent Internet · newsIntelligent Internet's Meta-Zenith Agent Rewrote vLLM Kernels for a 4x Speedup