APIUp to 25% cheaper than official pricesTry the API →
HermesHermes Agent Docs
Back to News

ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second

ByteShape ships ShapeLearn-quantized GGUFs of Qwen3.8-27B that fit on 12 GB GPUs and hit 176 tokens per second on an RTX 5090.

ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

ByteShape ships ShapeLearn-quantized GGUFs of Qwen3.8-27B that fit on 12 GB GPUs and hit 176 tokens per second on an RTX 5090.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report