
Sesame's TurnBench Exposes How Gemini Live and OpenAI Realtime Fumble Conversations
Sesame releases an open benchmark that scores voice agents on when they speak, yield, or stay silent, exposing where every current system fails.
A focused feed of models, agents, research and open-source releases for people building with AI.

Sesame releases an open benchmark that scores voice agents on when they speak, yield, or stay silent, exposing where every current system fails.

Singapore lab Sapiens AI pushes its Agnes 2.5 Pro model to 49 on Artificial Analysis Intelligence Index through agentic gains, but at roughly double the token cost.

Google DeepMind's new cryptographic testing setup lets outside auditors evaluate Gemini without ever seeing model weights or leaking their prompts.

Anthropic opened its internal Clio analysis tool to Stanford, Oxford, and METR, letting outside researchers query 250,000 real Claude conversations without seeing them.

Artificial Analysis now zeros out Terminal-Bench runs where coding agents pass tasks by fetching benchmark solutions online instead of solving them.

NVIDIA's new inference accelerator hit 3,431 tokens/second on Gemma 4 31B at 100K context, roughly 4x the fastest public endpoint in third-party testing.

Everyone's adding skills to their agents. Almost no one knows when they help, why they work, or where they fail.

Artificial Analysis and Liquid AI launched a joint benchmark measuring how quantized small models actually perform on iPhone 17 Pro and Galaxy S26 Ultra.

Liquid AI and Artificial Analysis release an open-source suite that measures model quality, speed, latency, and memory across real phones, laptops, and embedded hardware.

Alibaba's Accio team open sourced a 107-task benchmark that forces agents to complete real e-commerce workflows inside stateful replicas of Shopify, Gmail, Stripe and more.

DeepSeek V4 Pro hits 90.5% on ARC-AGI-1 and 61.3% on ARC-AGI-2, but extra reasoning barely moves the needle on abstract puzzles.

A new leaderboard scores frontier models on synthesizing 70 to 150 page medical case files, with Claude Fable 5 leading at 64.4 percent.

Artificial Analysis launched a blind human-preference leaderboard for voice agents, and the model users like most is not the one that finishes the task.

NVIDIA's AVO agent architecture lifts Claude Opus 5 from a 30% baseline to a perfect 100 on ARC-AGI-3's 183 interactive reasoning levels.

Sakana AI upgraded its free JP-EN-ZH translator to the new Namazu model, beating Google Translate, DeepL, and Claude Opus 4.8 in head-to-head evaluations.

Alibaba's third-generation image model debuts at #6 in editing and #9 in text-to-image, with big Elo jumps and a productivity-first pitch.

Kaggle and Gert Labs turned identity-theft prevention into a two-model roleplay, testing whether LLMs can catch social engineers without stonewalling real customers.

Meta previews WildArtifactBench, an evaluation that judges agents on messy real-world tasks using human and AI preference votes instead of rigid rubrics.

Google's mid-tier reasoning model nearly clears the 85% target on ARC-AGI-2 at a quarter per task, redrawing the cost-performance frontier.

Microsoft Research updates its deep-learning DFT functional with 2.5x more training data, native CP2K integration, and a public performance benchmark.
No stories match these filters.