←Back to NewsAI News/InfrarepoInfraGpusQwen3.8-27B Hits 417 tok/s on a Single RTX 3090 With 150k ContextA community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.SourceAlphaSignalPublishedAug 27, 2026, 8:51 AMAuthorAlphaSignal NewsroomRead1 min readA community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsTogether AI · newsTogether AI's ThunderKittens Hits 22.4 PFLOPS on NVIDIA's RubinIntelligent Internet · newsIntelligent Internet's Meta-Zenith Agent Rewrote vLLM Kernels for a 4x SpeedupPyTorch · newsNVIDIA's CUDA Python 1.0 Makes Python a First-Class GPU Citizen