
OpenAI's GPT-6 Astra Hits 99.9% on AGI Benchmarks and Runs Your Computer
OpenAI's new flagship saturates ARC-AGI-3 and FrontierMath Tier 4, drives a browser at 1.9x speed, and crosses the Critical cybersecurity threshold.
A focused feed of models, agents, research and open-source releases for people building with AI.

OpenAI's new flagship saturates ARC-AGI-3 and FrontierMath Tier 4, drives a browser at 1.9x speed, and crosses the Critical cybersecurity threshold.

Anthropic is proposing TypeScript Function Hooks for Claude Code, letting plugins deeply customize behavior with Express-style middleware and admin-controlled side effects.

Tencent Hunyuan's environment evolution builds off-policy lineages of increasingly hard terminal tasks, boosting Qwen3.6 agents by up to 18 points on Terminal-Bench.

Warp launches Factory Benchmarks, letting teams replay real agent runs to pick the best model per task and cut cost-per-PR from $80 to $30.

Google Research and HHMI Janelia released a wiring diagram of 166,000 neurons and 125 million synapses spanning the male fruit fly brain and nerve cord.

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

Keep the project on disk, treat lunch as a bill, and score the first finished tasks rather than dollars per million tokens

Runway's new general world model streams continuous 720p video and 48kHz audio in real time, taking text prompts and camera input as you play.

Browser Use agents can now check out on any website using disposable single-use cards issued by Stripe Link, without ever seeing your real card number.

Krea opens the beta for a new agent platform that turns natural-language prompts into finished visual assets, sitting on top of its existing creative suite.

Google DeepMind's new weather AI trains on raw satellite feeds and station data, delivering hourly 5km forecasts with up to 50% better rain accuracy.

Microsoft AI's new speech recognition model hits 2.0% word error rate at 411x real time speed for $1.67 per 1,000 minutes.

A new generative retriever leverages the table of contents to hit 82.6% Recall@1 on a fresh 18-book benchmark, with hallucinations under 0.05%.

Xiaohongshu's Self-GC uses a planner LLM to fold, mask, or prune agent context, retaining critical details 84.85% of the time versus 54.55% for rule-based baselines.

Figure locks in a $3.5B compute commitment scaling past $6B, targeting up to 100,000 Vera Rubin GPUs to train its Helix humanoid AI.

A new 365-day simulation drops LLM agents into a deterministic marketplace with ¥100,000, real supplier data, and fraudsters, and no model wins everything.

An 8-year interpretability project shows that LLM representations can be closely approximated by symbolic role-filler structures, enabling precise behavioral edits.

Alibaba's Zvec team open-sourced zg, a local-first CLI that fuses ripgrep, BM25, and vector search into one interface for humans and coding agents.

A ComfyUI custom node makes MiniMax H3 clips chain seamlessly by carrying motion and the exact audio waveform across cuts.

Cursor now lets teams run cloud coding agents on their own infrastructure or partnered sandbox providers, with auto-scaling worker pools.

Artificial Analysis rebuilt its Image Editing Arena with a two-axis taxonomy of editing actions and use cases, revealing which model wins each specific task.
Perplexity open-sources Lily, a Metal-based inference engine tuned for Qwen3.6-35B-A3B that beats MLX-LM by 1.23x prefill and 1.35x decode on M5 Max.

Meta's latest agentic model trims tool calls by 20 percent and tokens by 25 percent while learning when to ask for help.

Anthropic open-sources a full commerce agent blueprint with shopping and merchant reference implementations, plus architectural guidance from production deployments driving 35% larger carts.
No stories match these filters.