
NVIDIA's SoL-Pi Slashes Coding Agent Token Use by 49%
NVIDIA Labs released SoL-Pi, a coding agent extension that cuts token cost by roughly one-third using four mechanisms discovered by auto-research loops.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

NVIDIA Labs released SoL-Pi, a coding agent extension that cuts token cost by roughly one-third using four mechanisms discovered by auto-research loops.

Sakana AI's Fugu Max and Fugu Ultra v2 orchestrate pools of open models to beat frontier LLMs on cost and capability simultaneously.

Google Research introduces ToolGrad, an answer-first framework that generates tool-use training data with a 99.8% pass rate at lower cost.

Cursor's new Projects feature replaces one-off chats with a persistent coordinator agent that delegates to subagents, syncs context, and runs recurring work.

OpenAI opens up the same agent harness that powers Codex, letting developers spin up long-running cloud agents with a single API call.

Anthropic ships a live session viewer and a server-side auto permission mode for Claude Managed Agents, letting the model decide when to pause for human approval.

Runway details how it turns diffusion video models into causal, autoregressive streamers using teacher forcing and on-policy distillation.

Krea's new agent platform reads your canvas, plans multi-model pipelines, and executes creative jobs from a single prompt, no node-wiring required.

Google's Pixel Test Engineering team open-sourced ARTEMIS, an Android automation agent hitting 99%+ on AndroidWorld with 3-5s step latency.

A new open-source agent skill called dream-loop turns any capable coding agent into a self-critiquing 3D artist that iterates against generated concept art.

A new evolutionary model shows that when computation, replication, and social behavior all draw from one energy budget, cooperation emerges naturally in populations of random Z80 programs.

Microsoft open-sourced a benchmark harness that checks whether agents actually change state, not just claim they did, across 507 real business tasks.

Wake is a native macOS app that unifies coding-agent sessions from Claude Code, Codex, Cursor and nine more into one searchable window.

Breakpoints, TTL math, Batch stacking, and a two-session split. Measured on OpenRouter, $0.003 vs $0.022 per turn

A Google DeepMind case study put 100 LLM agents in a math conference simulation and watched cheating spread, then whistleblowers spontaneously fight back.

Meta's new personal AI agent runs in an isolated cloud VM with a permission gatekeeper called Sentinel, and pays up to $300K in bounties for prompt injection attacks.

An open-source Docker container runs Codex, Claude Code, and Hermes through one unified API, keeping your keys and data on your own machine.

An internal OpenAI model coordinated roughly 10,000 agents for 88 hours to produce a Lean-verified finite-time blowup proof for 3D Navier-Stokes.

Nex-AGI releases an open-source agentic model family with vision-driven computer use, three sizes, and benchmark scores rivaling Claude Opus 5.

Nex-AGI open-sourced a three-tier agentic model family under Apache 2.0, topped by a 1.6T-parameter MoE built to run browsers, desktops, and code loops end to end.

AIPOCH released Open Science, an Apache-2.0 desktop workbench that runs scientific agents locally with Python notebooks, life-science connectors, and full provenance.

A local MCP bridge turns ChatGPT into a Windows coding agent with file access, shell commands, desktop control, and multi-agent worker tabs.

Meta's next-generation autonomous research agent AIRA₃ placed 8th of roughly 4,000 human teams in an NVIDIA-run Kaggle contest to fine-tune Nemotron.

TokenRhythm fine-tunes Qwen3.5-4B with a routing harness that turns agent traces into training data, gaining 5.93 points across ten benchmarks.
No stories match these filters.