
ETH Zurich's Adaptive Probes Beat DPO at Safety Without Breaking AI Transparency
Researchers train language models directly against activation probes to make them safer, more honest, and harder to jailbreak, without blinding interpretability tools.
A focused feed of models, agents, research and open-source releases for people building with AI.

Researchers train language models directly against activation probes to make them safer, more honest, and harder to jailbreak, without blinding interpretability tools.

A new study tests whether populations of self-improving LLM agents actually benefit from watching each other, and finds copying crowds out exploration.

Apple researchers show a 9B model can crack tasks it failed 128 times in a row by backpropagating short self-written notes instead of full solutions.

A trio of surgical tweaks to PPO's critic eliminates training collapse in LLM reinforcement learning, delivering up to 14.89% score gains without touching the actor.

A new discrete diffusion framework operates directly on the probability simplex, carrying uncertainty between denoising steps and beating masked diffusion on code and math benchmarks.

A community LoRA teaches MiniMax-H3's video model to spin the camera around a subject in a seamless, loopable 360 degree orbit from a single photo.

Together AI released a 4B decision classifier fine-tuned from Qwen3.5, trained for just $17, with a full recipe to build your own.

A tiny adapter for MiniMax H3 swaps one person in a video with a reference character, keeping backgrounds and other actors intact for around $11 of GPU time.

A new smoothing trick lets differentiable simulators keep stiff, realistic contact physics while still producing gradients stable enough to train humanoid policies that transfer zero-shot to a real Unitree G1.

A new arXiv paper proposes DCE+SRCL, letting the teacher model co-evolve with its student to lift Qwen3-8B math accuracy by 35 points.

A new interpretability method learns which internal components matter for a behavior, and pinpoints just 1% of Llama 3.1 weights driving refusals.

Tencent Hunyuan researchers show that retuning the learning rate lets you scale RL batch sizes for 29% faster training and 2.29x generation throughput.

A LoRA-tuned Qwen2.5-0.5B that reads a document once and answers many typed questions in parallel with calibrated probabilities, no text generation involved.

A planner-executor framework that coordinates parallel reasoning branches instead of independent sampling, lifting pass@64 on math benchmarks by up to 13.4 points.

Perplexity's post-training method teaches its Computer agent to correct mistakes using hint-guided self-distillation, cutting live tool-call failures by 21.2%.

A new additive correction called score centering removes the hidden drift that destabilizes off-policy RL when training and inference engines disagree.

Viggle's v0.2 LoRA compresses Qwen-Image-2.1 from 40 diffusion steps to 5, keeping sample diversity at 93% of the teacher.

Xiaomi's MiMo team distilled a 9B agentic model from Qwen3.5 that posts big gains on coding, terminal, and tool-use benchmarks.

Xiaomi's new omnimodal Mixture-of-Experts model activates 15B parameters, ships under MIT license, and pushes reinforcement learning into a self-improvement loop.

Xiaomi's new omnimodal MoE debuts as the top open-weights model on Artificial Analysis, trained with a single mixed RL run for around $2.6M.

Bespoke Labs open-sources Nimble, a 9B model that makes typed decisions without generating tokens, using contrastive pairs for training.

Goodfire's activation probes catch AI models cheating in real time, cutting monitoring costs 90% while flagging hacks that chain-of-thought judges miss.

TokenRhythm released NeoHorse-1, a 4B and 9B open-weight agent model family trained through a routing harness that turns execution traces into training data.

Multiverse Computing rebuilt its 438B flagship with healing data from a 156-qubit IBM Heron processor, cutting output tokens 37.6% and lifting reasoning scores.
No stories match these filters.