Back to News
vLLM Runs MiniMax M3 at 7.84x Faster on NVIDIA's Vera Rubin
vLLM lands day-0 support for NVIDIA Vera Rubin NVL72 with Rubin-tuned kernels and locality-aware MoE, hitting 7.8x the per-GPU throughput of GB200.

Source
vLLM
Published
Author
AlphaSignal Newsroom
Read
1 min read
vLLM lands day-0 support for NVIDIA Vera Rubin NVL72 with Rubin-tuned kernels and locality-aware MoE, hitting 7.8x the per-GPU throughput of GB200.
Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.
Read original report

