←Back to NewsAI News/InfranewsInfraOpen SourceCohere's Open-Source Megakernel Beats vLLM by 1.58x on H100Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.SourceCoherePublishedSep 8, 2026, 7:44 PMAuthorAlphaSignal NewsroomRead1 min readCohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsPerplexity · newsPerplexity's CobbleDB Cuts Search Storage Latency by 82% Over DynamoDBAlphaSignal · repoEdge0 Runs a 35B AI Model on a Mac mini Using SSDvLLM · newsvLLM's Hybrid HiSparse Triples Concurrent Requests on Million-Token Contexts