hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Qwen3.8-27B Hits 417 tok/s on a Single RTX 3090 With 150k Context

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.

Qwen3.8-27B Hits 417 tok/s on a Single RTX 3090 With 150k Context
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report