hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

FlashAttention-4's Direct-P Finally Unlocks 2.13x FP4 Speedup on NVIDIA GB200

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

FlashAttention-4's Direct-P Finally Unlocks 2.13x FP4 Speedup on NVIDIA GB200
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report