←Back to NewsAI News/LlmsmodelLlmsGpusQwen Drops Qwen3.8-Flash-Next, a 125B Open Model That Activates Only 6B ParametersQwen ships an FP8 preview of the architecture behind Qwen4, pairing sparse attention, n-gram embeddings and 125B params with 6B active.SourceAlphaSignalPublishedAug 24, 2026, 8:25 AMAuthorAlphaSignal NewsroomRead1 min readQwen ships an FP8 preview of the architecture behind Qwen4, pairing sparse attention, n-gram embeddings and 125B params with 6B active.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsPrismML · newsPrismML Squeezes Qwen3.8 27B Into 5.9 GB With 98% Performance RetainedPrismML · modelPrism ML's Ternary Bonsai 2 Squeezes a 27B Reasoning Model Into 8.6 GBAlphaSignal · modelEmpero Distills Qwen3.8 Reasoning Into a Lean 35B Open Model