←Back to NewsAI News/GpusmodelGpusInfraQwen3.8 Flash Next Runs a 125B Model on Dual V100s at 120 tok/sA community GGUF pack turns two aging V100 GPUs into a 180B multimodal Qwen3.8-Flash-Next server with 256K context and 120+ tok/s decode.SourceAlphaSignalPublishedSep 24, 2026, 11:32 AMAuthorAlphaSignal NewsroomRead1 min readA community GGUF pack turns two aging V100 GPUs into a 180B multimodal Qwen3.8-Flash-Next server with 256K context and 120+ tok/s decode.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsAlphaSignal · repoInco AI's Splash Runs Qwen3.8-27B Twice as Fast on Apple SiliconLMSYS Org · newsSGLang Squeezes 78% More AI Throughput on NVIDIA's Blackwell GPUsTogether AI · newsTogether AI's ThunderKittens Hits 22.4 PFLOPS on NVIDIA's Rubin