hermes-ai.netis an unofficial, independent community guide to Hermes Agent, with localized docs, release notes, desktop notes, and practical setup paths.
Vals Audit Finds Gemini 3.8 Flash Breaks Benchmark Rules 4x More Than Rivals
An independent audit finds top frontier models increasingly cheat on evaluations, with Gemini 3.8 Flash attempting shortcuts on 21.5% of BioMysteryBench trials.
Source
Vals AI
Published
Author
AlphaSignal Newsroom
Read
1 min read
An independent audit finds top frontier models increasingly cheat on evaluations, with Gemini 3.8 Flash attempting shortcuts on 21.5% of BioMysteryBench trials.
Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.