hermes-ai.netis an unofficial, independent community guide to Hermes Agent, with localized docs, release notes, desktop notes, and practical setup paths.
FlashAttention-4's Direct-P Finally Unlocks 2.13x FP4 Speedup on NVIDIA GB200
A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read
A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.
Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.