←Back to NewsAI News/LlmsnewsLlmsApiInception's Mercury 2.5 Hits 1,107 Tokens per Second, Beating Autoregressive ModelsInception's new diffusion-based LLM hits 1,107 tokens per second on NVIDIA GPUs while boosting quality 40% over Mercury 2.SourceInceptionPublishedSep 8, 2026, 4:45 PMAuthorAlphaSignal NewsroomRead1 min readInception's new diffusion-based LLM hits 1,107 tokens per second on NVIDIA GPUs while boosting quality 40% over Mercury 2.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsOpenAI · newsOpenAI's Astra for Law Pushes Into Big Firms With 54% Research AccuracySakana AI · newsSakana AI's Fugu Max Beats Frontier Models at 60% Lower CostGoogle DeepMind · newsGoogle's Gemini 3.8 Live Adds Real-Time Reasoning to Voice AI