
Alibaba's Qwen3.8-Flash-Next Shrinks 125B Model to 84 GiB on One Machine
A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.

A 26B multimodal Gemma variant with refusal directions surgically removed lands on Hugging Face in GGUF format, ready for llama.cpp.

A community fine-tune of Qwen3.8-27B strips refusals, fixes the resulting weight damage, and ships in four formats tuned for local inference.

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.

Red Hat AI shipped an NVFP4 quantization of Z.ai's 320B GLM-5.3-Flash, shrinking the model to run on vLLM with FP4 activations while holding reasoning benchmarks near the original.

The latest vLLM release lands 584 commits from 270 contributors, with big performance wins for Kimi-K3, DeepSeek-V4, and speculative decoding.

Cerebras unveiled the CS-4 wafer-scale system architecture at Hot Chips and sketched a roadmap to CS-5 and 3D-stacked DRAM in CS-6.

NVIDIA Dynamo's new shadow engine keeps a warm standby ready on the same GPU, cutting LLM failover from minutes to seconds.

Apple officially endorsed exo on its Mac Studio and Mac Mini pages, blessing a workflow that turns four desktops into a 4.8TB/s inference rig.

OpenAI's first in-house inference chip delivers 1.5-1.9x more work per watt and up to 4.1x lower latency versus current systems.

NVIDIA's new inference accelerator hit 3,431 tokens/second on Gemma 4 31B at 100K context, roughly 4x the fastest public endpoint in third-party testing.

Liquid AI and Artificial Analysis release an open-source suite that measures model quality, speed, latency, and memory across real phones, laptops, and embedded hardware.

Qwen ships an FP8 preview of the architecture behind Qwen4, pairing sparse attention, n-gram embeddings and 125B params with 6B active.

A community quantization pairs a 27B Qwen model with multi-token prediction and AMD's IU4 matrix path, hitting ~49 tokens per second on a single Strix Halo APU.

A community-made checkpoint grafts Z-Image's texture attention onto MiniMax H3's video engine, giving richer surfaces without retraining or extra VRAM.

A community-built vLLM container brings NVFP4 KV cache, DFlash speculative decoding, and Blackwell sm_121a runtime patches to DGX Spark serving.

A per-tensor dynamic quantization of an abliterated Qwen3.8-27B lands on Hugging Face, delivering 82.98% MMLU and 262k context on a single RTX 4090.
No stories match these filters.