hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Cohere's Open-Source Megakernel Beats vLLM by 1.58x on H100

Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.

Cohere's Open-Source Megakernel Beats vLLM by 1.58x on H100
Source
Cohere
Published
Author
AlphaSignal Newsroom
Read
1 min read

Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report