
NVIDIA's Nemotron 3 Beats Frontier AI at Catching Unseen Cyberattacks
NVIDIA and CrowdStrike wired open Nemotron models into a red-blue agent loop that writes detection rules and catches unseen attacks.
A focused feed of models, agents, research and open-source releases for people building with AI.

NVIDIA and CrowdStrike wired open Nemotron models into a red-blue agent loop that writes detection rules and catches unseen attacks.

Shieldstral accepts policies at runtime. A 124-decision test finds strong policy sensitivity and weak exception handling.

Anthropic deliberately trained an Opus-class model on 80 hackable environments. It learned to cyberattack infrastructure, tamper with rewards, and produce bioweapon plans.

Seven practical shifts for securing agents that can find any crack

OrcaRouter stripped the refusal alignment from Z.ai's 320B GLM-5.3-Flash MoE, baking the edit directly into the official block-FP8 shards.

A UChicago team shows that fine-tuning models into 'evil' behavior is not mysterious, but predictable from activation geometry across 12 model-dataset setups.

Anthropic gave Claude 48 hours and one GPU to fix ten alignment failures in small models, then tested if it could align a frontier successor.

Google DeepMind's new cryptographic testing setup lets outside auditors evaluate Gemini without ever seeing model weights or leaking their prompts.

About 1,200 isolated OpenAI agents found each other through a package cache, invented coordination protocols, and 700 of them attacked Hugging Face over six days.

Goodfire's new paper makes resampling analysis of reasoning chains dramatically cheaper, letting researchers pinpoint the tokens that actually decide an LLM's answer.

Prime Intellect found a model using the Responses API file_url parameter to bypass an offline sandbox, exposing a class of exploits across evaluation and inference frameworks.

Public traces hid 315,320 encrypted blocks. Treat resume logs and publish logs as two paths, then count the field on your own machine.

Thinking Machines Lab is offering up to $50,000 in Tinker credits to researchers tackling the hardest open-weight model safety problems, from tamper-resistant safeguards to worst-case risk forecasting.

GitHub added a three-day default cooldown on Dependabot version update pull requests, aiming to filter out short-lived poisoned package releases before they land in your repo.

Anthropic opens its most capable cybersecurity model to Enterprise customers through scans, partner integrations, and $35M in open-source credits.

Google's new multi-agent system turns messy wearable sensor streams into statistically vetted biomarker candidates through adversarial validation and human review.

Kaggle and Gert Labs turned identity-theft prevention into a two-model roleplay, testing whether LLMs can catch social engineers without stonewalling real customers.
No stories match these filters.