←Back to NewsAI News/GpusnewsGpusInfraTogether AI's ThunderKittens Hits 22.4 PFLOPS on NVIDIA's RubinTogether AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.SourceTogether AIPublishedSep 10, 2026, 6:00 PMAuthorAlphaSignal NewsroomRead1 min readTogether AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsIntelligent Internet · newsIntelligent Internet's Meta-Zenith Agent Rewrote vLLM Kernels for a 4x SpeedupPyTorch · newsNVIDIA's CUDA Python 1.0 Makes Python a First-Class GPU CitizenAlphaSignal · paperFlashAttention-4's Direct-P Finally Unlocks 2.13x FP4 Speedup on NVIDIA GB200