
AI agent skills: Why they succeed and what causes them to crash
Everyone's adding skills to their agents. Almost no one knows when they help, why they work, or where they fail.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Everyone's adding skills to their agents. Almost no one knows when they help, why they work, or where they fail.

Artificial Analysis and Liquid AI launched a joint benchmark measuring how quantized small models actually perform on iPhone 17 Pro and Galaxy S26 Ultra.

Liquid AI and Artificial Analysis release an open-source suite that measures model quality, speed, latency, and memory across real phones, laptops, and embedded hardware.

Alibaba's Accio team open sourced a 107-task benchmark that forces agents to complete real e-commerce workflows inside stateful replicas of Shopify, Gmail, Stripe and more.

DeepSeek V4 Pro hits 90.5% on ARC-AGI-1 and 61.3% on ARC-AGI-2, but extra reasoning barely moves the needle on abstract puzzles.

A new leaderboard scores frontier models on synthesizing 70 to 150 page medical case files, with Claude Fable 5 leading at 64.4 percent.

Artificial Analysis launched a blind human-preference leaderboard for voice agents, and the model users like most is not the one that finishes the task.

NVIDIA's AVO agent architecture lifts Claude Opus 5 from a 30% baseline to a perfect 100 on ARC-AGI-3's 183 interactive reasoning levels.

Sakana AI upgraded its free JP-EN-ZH translator to the new Namazu model, beating Google Translate, DeepL, and Claude Opus 4.8 in head-to-head evaluations.

Alibaba's third-generation image model debuts at #6 in editing and #9 in text-to-image, with big Elo jumps and a productivity-first pitch.

Kaggle and Gert Labs turned identity-theft prevention into a two-model roleplay, testing whether LLMs can catch social engineers without stonewalling real customers.

Meta previews WildArtifactBench, an evaluation that judges agents on messy real-world tasks using human and AI preference votes instead of rigid rubrics.

Google's mid-tier reasoning model nearly clears the 85% target on ARC-AGI-2 at a quarter per task, redrawing the cost-performance frontier.

Microsoft Research updates its deep-learning DFT functional with 2.5x more training data, native CP2K integration, and a public performance benchmark.
No stories match these filters.