
GPT-6 Astra Cracks a Decade-Old Voting Theory Problem Nobody Could Solve
Epoch AI marks the first Major Advance on its unsolved-math benchmark, with GPT-6 Astra driving the proof in an interactive session with three mathematicians.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Epoch AI marks the first Major Advance on its unsolved-math benchmark, with GPT-6 Astra driving the proof in an interactive session with three mathematicians.

Sakana AI launches the Frontier Intelligence Group, a research collective betting that Transformers and scaling laws aren't the final answer to AGI.

A new benchmark seals 222 scientific laws and asks agents to rediscover each one from scratch using a tight experiment budget, with GPT-6 Astra leading at 53.2%.

A new architecture proposes latent reasoning that grows with sequence length, sharing one recurrent state across prompt, response, training, and RL replay.

A math benchmark built to resist AI just fell. GPT-6 Astra cracked the final Tier 4 problem, closing out a 98 percent run in 14 months.

A new recurrent reasoning method mixes denoising with looped hidden states, hitting 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 with just 7M parameters.

Together AI highlights that Moonshot's open-weight Kimi K3 outscores Anthropic's newer Claude Fable 5.1 by 60% on the hard slice of Harvey's autonomous legal agent benchmark.

A new evolutionary model shows that when computation, replication, and social behavior all draw from one energy budget, cooperation emerges naturally in populations of random Z80 programs.

A new paper shows recurrent-depth reasoning models behave like chaotic dynamical systems, where hard problems create fractal basins that trap thinking near wrong answers.

An internal OpenAI model coordinated roughly 10,000 agents for 88 hours to produce a Lean-verified finite-time blowup proof for 3D Navier-Stokes.

A new physics inspired bound explains which patterns stochastic gradient descent learns first, tying acquisition speed to Fisher information flow.

Meta's next-generation autonomous research agent AIRA₃ placed 8th of roughly 4,000 human teams in an NVIDIA-run Kaggle contest to fine-tune Nemotron.

Claude autonomously wrote a 13 million line Lean proof of Fermat's Last Theorem in 11 days, verifying 29,500 supporting theorems along the way.

An 8-year interpretability project shows that LLM representations can be closely approximated by symbolic role-filler structures, enabling precise behavioral edits.

MIT researchers put hundreds of identical LLM agents into a persistent world and watched them evolve specialization, tool inheritance, and technology that survives without them.

A UIUC and Bridgewater team fine-tuned Kimi-K2.6 on Tinker to become the first text-to-SQL model to beat human accuracy.

A new method throws away most of a model's own reasoning trace mid-thought, cutting memory to a fixed cap and running inference 3x faster.

Goodfire's new paper makes resampling analysis of reasoning chains dramatically cheaper, letting researchers pinpoint the tokens that actually decide an LLM's answer.

Jeremy Avigad argues the fixation on neural theorem provers hides a far richer landscape of ways AI is reshaping how mathematics gets done.

DeepSeek V4 Pro hits 90.5% on ARC-AGI-1 and 61.3% on ARC-AGI-2, but extra reasoning barely moves the needle on abstract puzzles.
No stories match these filters.