
Prime Intellect's Prime Agent Hits 95.5% on ARC-AGI-3 by Rebuilding AI Scaffolding
Prime Intellect's open-source harness turns agents into programmable systems, driving ARC-AGI-3 scores from 30% to 95.5% and enabling week-long autonomous runs.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Prime Intellect's open-source harness turns agents into programmable systems, driving ARC-AGI-3 scores from 30% to 95.5% and enabling week-long autonomous runs.

Everyone's adding skills to their agents. Almost no one knows when they help, why they work, or where they fail.

A new Zig-built plugin turns x64dbg into an MCP endpoint, letting Claude and other AI agents drive reverse engineering sessions through 71 tools.

Alibaba's Accio team open sourced a 107-task benchmark that forces agents to complete real e-commerce workflows inside stateful replicas of Shopify, Gmail, Stripe and more.

GooeyPi wraps three terminal coding agents in a single Electron desktop app, giving Pi, OMP, and Prime Agent a shared cross platform GUI.

Rakazo bundles persistent AI teammates with their own live Linux desktops into a self-hosted stack you run on your own hardware and model keys.

TrueFoundry open-sourced TrueForge, an MIT-licensed agent runtime that ran 30-75% cheaper than Claude Managed Agents on a 14-task enterprise benchmark.

NVIDIA's AVO agent architecture lifts Claude Opus 5 from a 30% baseline to a perfect 100 on ARC-AGI-3's 183 interactive reasoning levels.

DeepSeek open-sourced dsh, a plugin-based coding agent framework where the model, tools, sandbox, UI, and even the agent loop itself are all swappable.

An unofficial Tauri 2 desktop app wraps xAI's Grok Build CLI with sessions, project management, media previews, and scheduled automations.

A Swift CLI from LY Corporation gives AI coding agents a token-efficient observe-act loop across iOS Simulator and Android devices.

Magnitude's new catalog profiles your hardware, estimates tokens per second for every model, and picks the best local LLM before you download anything.

Anthropic makes computer use, browser tool, Skills API, and Files API generally available with batched actions cutting round trips 20-40%.

Meta previews WildArtifactBench, an evaluation that judges agents on messy real-world tasks using human and AI preference votes instead of rigid rubrics.

AT&T takes an equity stake in Brett Adcock's stealthy AI hardware startup Hark, providing cellular connectivity for phone-free AI devices.

AWS ships a governed bridge between coding agents and 300+ services, replacing AWS Labs tooling with plugins, skills, and a managed MCP server.

AMAP-ML released an open-source orchestration layer that lets Claude Code, Codex, and OpenClaw agents complete multi-hour computer tasks without state drift.
No stories match these filters.