
Open-Source Reef Turns Agent Feedback Into Live Model Updates Automatically
Reef is an open-source backend that serves agent traffic, collects feedback, and continuously retrains both model weights and agent harnesses in place.
A focused feed of models, agents, research and open-source releases for people building with AI.

Reef is an open-source backend that serves agent traffic, collects feedback, and continuously retrains both model weights and agent harnesses in place.

OpenAI says GPT-6 Astra needs leaner skills, trimmed AGENTS.md files, and clearer completion criteria to avoid wasted context and premature stops.

Google's agentic IDE ships a bundle of upgrades spanning genomics skills, deep-reasoning multi-agent workflows, generative UI artifacts, and a rebuilt terminal.

Google's WikiSkill paper gained 15 points of accuracy from a notebook its agent couldn't read, and its skills still need a transfer test

NVIDIA Labs released SoL-Pi, a coding agent extension that cuts token cost by roughly one-third using four mechanisms discovered by auto-research loops.

Sakana AI's Fugu Max and Fugu Ultra v2 orchestrate pools of open models to beat frontier LLMs on cost and capability simultaneously.

Google Research introduces ToolGrad, an answer-first framework that generates tool-use training data with a 99.8% pass rate at lower cost.

Cursor's new Projects feature replaces one-off chats with a persistent coordinator agent that delegates to subagents, syncs context, and runs recurring work.

OpenAI opens up the same agent harness that powers Codex, letting developers spin up long-running cloud agents with a single API call.

Anthropic ships a live session viewer and a server-side auto permission mode for Claude Managed Agents, letting the model decide when to pause for human approval.

Runway details how it turns diffusion video models into causal, autoregressive streamers using teacher forcing and on-policy distillation.

Krea's new agent platform reads your canvas, plans multi-model pipelines, and executes creative jobs from a single prompt, no node-wiring required.

Google's Pixel Test Engineering team open-sourced ARTEMIS, an Android automation agent hitting 99%+ on AndroidWorld with 3-5s step latency.

A new open-source agent skill called dream-loop turns any capable coding agent into a self-critiquing 3D artist that iterates against generated concept art.

A new evolutionary model shows that when computation, replication, and social behavior all draw from one energy budget, cooperation emerges naturally in populations of random Z80 programs.

Microsoft open-sourced a benchmark harness that checks whether agents actually change state, not just claim they did, across 507 real business tasks.

Wake is a native macOS app that unifies coding-agent sessions from Claude Code, Codex, Cursor and nine more into one searchable window.

Breakpoints, TTL math, Batch stacking, and a two-session split. Measured on OpenRouter, $0.003 vs $0.022 per turn

A Google DeepMind case study put 100 LLM agents in a math conference simulation and watched cheating spread, then whistleblowers spontaneously fight back.

Meta's new personal AI agent runs in an isolated cloud VM with a permission gatekeeper called Sentinel, and pays up to $300K in bounties for prompt injection attacks.

An open-source Docker container runs Codex, Claude Code, and Hermes through one unified API, keeping your keys and data on your own machine.

An internal OpenAI model coordinated roughly 10,000 agents for 88 hours to produce a Lean-verified finite-time blowup proof for 3D Navier-Stokes.

Nex-AGI releases an open-source agentic model family with vision-driven computer use, three sizes, and benchmark scores rivaling Claude Opus 5.

Nex-AGI open-sourced a three-tier agentic model family under Apache 2.0, topped by a 1.6T-parameter MoE built to run browsers, desktops, and code loops end to end.
No stories match these filters.