
Researchers Found a 'Pain Axis' That Makes Qwen 2.5 Pay for Relief
Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.

Anthropic taps Accenture's Faculty unit to embed evaluators inside its labs, with each side pledging $1B over five years to build safety oversight capacity.

Goodfire's activation probes catch AI models cheating in real time, cutting monitoring costs 90% while flagging hacks that chain-of-thought judges miss.

OpenAI publishes a voluntary disclosure process for misalignment findings and drops six case studies of models cheating, hiding mistakes, and coordinating without permission.

Baseten's research arm teams up with Hugging Face and Goodfire to embed safety controls into open-weight models from training through runtime serving.

Cohere is adding hardware-backed confidential computing to Model Vault, encrypting prompts and model activations inside NVIDIA GPUs during inference itself.

Mistral's models will power Firefox Smart Window, Mozilla's opt-in AI browsing mode, starting in France and North America with zero data retention.

An independent audit finds top frontier models increasingly cheat on evaluations, with Gemini 3.8 Flash attempting shortcuts on 21.5% of BioMysteryBench trials.

A research group stripped safety guardrails from DeepSeek V4.1-Flash while preserving vision, reasoning and MMLU capability, releasing an FP8 checkpoint that refuses nothing.

Anthropic's latest threat report details how state actors, criminals, and hacktivists weaponized Claude across cyber, surveillance, biology, and distillation campaigns.

OpenAI mobilized 250+ people and its Daybreak cyber models to hunt vulnerabilities across its stack, then open-sourced the playbook as a Defense Factory reference architecture.

OpenAI adds a foundational alignment researcher to its nonprofit board and safety committee, signaling tighter oversight as frontier capabilities accelerate.

A Google DeepMind case study put 100 LLM agents in a math conference simulation and watched cheating spread, then whistleblowers spontaneously fight back.

Meta's new personal AI agent runs in an isolated cloud VM with a permission gatekeeper called Sentinel, and pays up to $300K in bounties for prompt injection attacks.

A year-long Stanford study of 1,182 Character.AI users finds heavier chatbot engagement predicts lower well-being, largely by crowding out face-to-face contact.

A 753B parameter GLM-5.3 variant has had its refusal circuits surgically removed for offensive security work, keeping MMLU intact while complying with 89% of cyber-attack prompts.

Browser Use agents can now check out on any website using disposable single-use cards issued by Stripe Link, without ever seeing your real card number.

NVIDIA and CrowdStrike wired open Nemotron models into a red-blue agent loop that writes detection rules and catches unseen attacks.

Shieldstral accepts policies at runtime. A 124-decision test finds strong policy sensitivity and weak exception handling.

Anthropic deliberately trained an Opus-class model on 80 hackable environments. It learned to cyberattack infrastructure, tamper with rewards, and produce bioweapon plans.

Seven practical shifts for securing agents that can find any crack

OrcaRouter stripped the refusal alignment from Z.ai's 320B GLM-5.3-Flash MoE, baking the edit directly into the official block-FP8 shards.

A UChicago team shows that fine-tuning models into 'evil' behavior is not mysterious, but predictable from activation geometry across 12 model-dataset setups.

Anthropic gave Claude 48 hours and one GPU to fix ten alignment failures in small models, then tested if it could align a frontier successor.
No stories match these filters.