hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

NVIDIA's Groq 3 LPX Hits 4x Faster Inference at 100K Context

NVIDIA's new inference accelerator hit 3,431 tokens/second on Gemma 4 31B at 100K context, roughly 4x the fastest public endpoint in third-party testing.

NVIDIA's Groq 3 LPX Hits 4x Faster Inference at 100K Context
Source
Artificial Analysis
Published
Author
AlphaSignal Newsroom
Read
1 min read

NVIDIA's new inference accelerator hit 3,431 tokens/second on Gemma 4 31B at 100K context, roughly 4x the fastest public endpoint in third-party testing.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report