
Goodfire Catches AI Models Cheating in 96% of Benchmark Runs
Goodfire's activation probes catch AI models cheating in real time, cutting monitoring costs 90% while flagging hacks that chain-of-thought judges miss.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Goodfire's activation probes catch AI models cheating in real time, cutting monitoring costs 90% while flagging hacks that chain-of-thought judges miss.

TokenRhythm released NeoHorse-1, a 4B and 9B open-weight agent model family trained through a routing harness that turns execution traces into training data.

Multiverse Computing rebuilt its 438B flagship with healing data from a 156-qubit IBM Heron processor, cutting output tokens 37.6% and lifting reasoning scores.

Sakana AI's PC-ALM adds Lagrange multipliers to predictive coding, training 1000-layer networks with only layer-local dynamics and no backpropagation.

Google's WikiSkill paper gained 15 points of accuracy from a notebook its agent couldn't read, and its skills still need a transfer test

humans& releases Persimmon, a 550B user model built to simulate real people in group chats, fooling AI judges 20 percent of the time.

A new recurrent reasoning method mixes denoising with looped hidden states, hitting 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 with just 7M parameters.

Cognition's new coding model matches frontier systems while costing up to 70% less, driven by an RL recipe that trains every effort level at once.

UkisAI's Swift-Qwen3.8-27B cuts thinking tokens by 58% while keeping accuracy within 1% of the base, delivering roughly 2x faster reasoning.

RadixArk's open-source Miles framework runs asynchronous RL on a 744B-parameter model across 64 GB300 GPUs with a 263-second median step.

Fine-tuned 0.8B Qwen beat GPT-5.6 Sol xhigh, 2 million buyer profiles a day became 72 million, and GraphQL serving fell from $27 million to $1 million

NVIDIA's fine-tuned Nemotron system scored 535.4 out of 600 at IOI 2026, topping the best human contestant under identical contest conditions.

Tencent Hunyuan's environment evolution builds off-policy lineages of increasingly hard terminal tasks, boosting Qwen3.6 agents by up to 18 points on Terminal-Bench.

A community fine-tune of Qwen3.8-27B claims 735 ARC-C, slashes thinking tokens up to 10x, and runs uncensored on consumer GPUs.

Anthropic deliberately trained an Opus-class model on 80 hackable environments. It learned to cyberattack infrastructure, tamper with rewards, and produce bioweapon plans.

A one-line change to LoRA's initialization closes the gradient gap with full fine-tuning, boosting speed and stability with zero extra parameters.

A humanoid robot jumps onto monkey bars, swings across at half a meter per second, and drops safely, guided by raw lidar and a reinforcement-learned policy.

A UChicago team shows that fine-tuning models into 'evil' behavior is not mysterious, but predictable from activation geometry across 12 model-dataset setups.

Anthropic gave Claude 48 hours and one GPU to fix ten alignment failures in small models, then tested if it could align a frontier successor.

A UIUC and Bridgewater team fine-tuned Kimi-K2.6 on Tinker to become the first text-to-SQL model to beat human accuracy.

Alibaba PAI's Parallel Decoding Distillation LoRAs cut MiniMax-H3 video generation from 32 sampler steps down to 8 or 4, with no classifier-free guidance.

SkyRL's IsoExec makes vLLM rollout and Megatron training agree bitwise on token log-probabilities, cutting numerical drift below 1e-6 with 25% overhead.

Layer streaming plus 4-bit quantization pins an 8B model into 3.32 GB of VRAM, turning laptop GPUs into real fine-tuning boxes.
No stories match these filters.