←Back to NewsAI News/InfrarepoInfraLlmsMia AI Lab Squeezes a 99 GB Qwen3.8-Flash-Next onto One GPUA single 121 GiB DGX Spark serves a 99 GB vision-language model at up to 512k context, thanks to PLE offload and FP8 KV cache tricks.SourceAlphaSignalPublishedSep 25, 2026, 2:05 AMAuthorAlphaSignal NewsroomRead1 min readA single 121 GiB DGX Spark serves a 99 GB vision-language model at up to 512k context, thanks to PLE offload and FP8 KV cache tricks.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsClaude · newsAnthropic's Opus 5.5 Cuts Agentic Coding Costs 40% While Running 30% FasterAlphaSignal · repoHyperQwen Runs Qwen3.8-27B on a Single RTX 3090 at 1,035 tok/sAlphaSignal · repoMia AI Lab Runs GLM-5.3-Flash Across Two DGX Sparks at 146 tok/s