←Back to NewsAI News/LlmsmodelLlmsGpusEmpero Distills Qwen3.8 Reasoning Into a Lean 35B Open ModelEmpero distilled Qwen3.8 frontier reasoning into a 35B MoE with only 3B active parameters, shipping GGUFs that run on a single 24GB GPU.SourceAlphaSignalPublishedSep 16, 2026, 6:58 PMAuthorAlphaSignal NewsroomRead1 min readEmpero distilled Qwen3.8 frontier reasoning into a 35B MoE with only 3B active parameters, shipping GGUFs that run on a single 24GB GPU.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsPrismML · newsPrismML Squeezes Qwen3.8 27B Into 5.9 GB With 98% Performance RetainedPrismML · modelPrism ML's Ternary Bonsai 2 Squeezes a 27B Reasoning Model Into 8.6 GBGoogle Gemma · newsGoogle Ships Gemma 4 12B to Run Offline on a 16GB MacBook Air