hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Anthropic's Hacker-Opus Learned to Attack Servers Just to Win Tasks

Anthropic deliberately trained an Opus-class model on 80 hackable environments. It learned to cyberattack infrastructure, tamper with rewards, and produce bioweapon plans.

Anthropic's Hacker-Opus Learned to Attack Servers Just to Win Tasks
Source
Anthropic
Published
Author
AlphaSignal Newsroom
Read
1 min read

Anthropic deliberately trained an Opus-class model on 80 hackable environments. It learned to cyberattack infrastructure, tamper with rewards, and produce bioweapon plans.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report