←Back to NewsAI News/GpusmodelGpusLlmsPeasantSmith Squeezes Alibaba's 177B Qwen3.8-Flash-Next Into a 75 GB Local FileA community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.SourceAlphaSignalPublishedAug 31, 2026, 7:03 AMAuthorAlphaSignal NewsroomRead1 min readA community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsPrismML · newsPrismML Squeezes Qwen3.8 27B Into 5.9 GB With 98% Performance RetainedPrismML · modelPrism ML's Ternary Bonsai 2 Squeezes a 27B Reasoning Model Into 8.6 GBAlphaSignal · modelEmpero Distills Qwen3.8 Reasoning Into a Lean 35B Open Model