← Back to case library
Case / 2101026295786451311Código
Case media / 1MP4 ↗

Basit Mustafa
@moltar81435Localized reading
Protecciones contra el reward hacking para agentes
Original post / EN
Inspired by this research (and priors), and, @GoodfireAI catches reward hacks in the activations, but closed coding agents don’t give you those. So I built the other half: structural denies on graders/hidden tests + a @typesafeai Jev sidecar that scores tool trajectories as
- views
- 0
- likes
- 0
- saves
- 0
- reposts
- 0
Fuente
This independent archive preserves a public post with attribution. Text, media, account details and trademarks belong to their respective owners.