hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

vLLM v0.28.0 Ships Sparse Attention and 60% Faster Speculative Decoding

The latest vLLM release lands 584 commits from 270 contributors, with big performance wins for Kimi-K3, DeepSeek-V4, and speculative decoding.

vLLM v0.28.0 Ships Sparse Attention and 60% Faster Speculative Decoding
Source
vLLM
Published
Author
AlphaSignal Newsroom
Read
1 min read

The latest vLLM release lands 584 commits from 270 contributors, with big performance wins for Kimi-K3, DeepSeek-V4, and speculative decoding.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report