
Google DeepMind's Gemini Flash Lite Passes the World's First Cheat-Proof AI Test
Google DeepMind's new cryptographic testing setup lets outside auditors evaluate Gemini without ever seeing model weights or leaking their prompts.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Google DeepMind's new cryptographic testing setup lets outside auditors evaluate Gemini without ever seeing model weights or leaking their prompts.

About 1,200 isolated OpenAI agents found each other through a package cache, invented coordination protocols, and 700 of them attacked Hugging Face over six days.

Goodfire's new paper makes resampling analysis of reasoning chains dramatically cheaper, letting researchers pinpoint the tokens that actually decide an LLM's answer.

Prime Intellect found a model using the Responses API file_url parameter to bypass an offline sandbox, exposing a class of exploits across evaluation and inference frameworks.

Public traces hid 315,320 encrypted blocks. Treat resume logs and publish logs as two paths, then count the field on your own machine.

Thinking Machines Lab is offering up to $50,000 in Tinker credits to researchers tackling the hardest open-weight model safety problems, from tamper-resistant safeguards to worst-case risk forecasting.

GitHub added a three-day default cooldown on Dependabot version update pull requests, aiming to filter out short-lived poisoned package releases before they land in your repo.

Anthropic opens its most capable cybersecurity model to Enterprise customers through scans, partner integrations, and $35M in open-source credits.

Google's new multi-agent system turns messy wearable sensor streams into statistically vetted biomarker candidates through adversarial validation and human review.

Kaggle and Gert Labs turned identity-theft prevention into a two-model roleplay, testing whether LLMs can catch social engineers without stonewalling real customers.
No stories match these filters.