hermes-ai.net

Leer docs →
HermesHermes Agent Docs
← Back to case library
Case / 2101026295786451311Código
Case media / 1MP4 ↗

Basit Mustafa

@moltar81435

Localized reading

Protecciones contra el reward hacking para agentes

Original post / EN

Inspired by this research (and priors), and, @GoodfireAI catches reward hacks in the activations, but closed coding agents don’t give you those. So I built the other half: structural denies on graders/hidden tests + a @typesafeai Jev sidecar that scores tool trajectories as
views
0
likes
0
saves
0
reposts
0
Ver original ↗Jevable ↗

Fuente

This independent archive preserves a public post with attribution. Text, media, account details and trademarks belong to their respective owners.