hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Alibaba's Qwen3.8-Flash-Next Shrinks 125B Model to 84 GiB on One Machine

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

Alibaba's Qwen3.8-Flash-Next Shrinks 125B Model to 84 GiB on One Machine
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report