
Anthropic's Claude Formally Proved Fermat's Last Theorem in 11 Days
Claude autonomously wrote a 13 million line Lean proof of Fermat's Last Theorem in 11 days, verifying 29,500 supporting theorems along the way.
A focused feed of models, agents, research and open-source releases for people building with AI.

Claude autonomously wrote a 13 million line Lean proof of Fermat's Last Theorem in 11 days, verifying 29,500 supporting theorems along the way.

An 8-year interpretability project shows that LLM representations can be closely approximated by symbolic role-filler structures, enabling precise behavioral edits.

MIT researchers put hundreds of identical LLM agents into a persistent world and watched them evolve specialization, tool inheritance, and technology that survives without them.

A UIUC and Bridgewater team fine-tuned Kimi-K2.6 on Tinker to become the first text-to-SQL model to beat human accuracy.

A new method throws away most of a model's own reasoning trace mid-thought, cutting memory to a fixed cap and running inference 3x faster.

Goodfire's new paper makes resampling analysis of reasoning chains dramatically cheaper, letting researchers pinpoint the tokens that actually decide an LLM's answer.

Jeremy Avigad argues the fixation on neural theorem provers hides a far richer landscape of ways AI is reshaping how mathematics gets done.

DeepSeek V4 Pro hits 90.5% on ARC-AGI-1 and 61.3% on ARC-AGI-2, but extra reasoning barely moves the needle on abstract puzzles.
No stories match these filters.