
Vals AI's MysteryMechanism Benchmark Shows GPT-6 Astra Tops Science Test at 53%
A new Vals AI benchmark forces frontier models to rediscover hidden math laws through experiments, exposing where extra reasoning stops paying off.
A focused feed of models, agents, research and open-source releases for people building with AI.

A new Vals AI benchmark forces frontier models to rediscover hidden math laws through experiments, exposing where extra reasoning stops paying off.

A Google team shows tiny grids of locally-connected cells can solve Sudoku, mazes, and ARC-AGI through emergent, visible reasoning dynamics.

A new study of 270 puzzles with multiple correct answers finds AI models cluster tightly together and diverge sharply from how humans pick solutions.

Sakana AI and the University of Tokyo show that a vision language model can plan robot motions more reliably by testing and refining candidate trajectories in simulation before touching the real robot.

Anthropic's Claude autonomously pushed a frontier theoretical physics calculation from eight to nine loops on Claude Science, costing only a few thousand dollars.

A new arXiv paper proposes DCE+SRCL, letting the teacher model co-evolve with its student to lift Qwen3-8B math accuracy by 35 points.

Vals AI ran GPT-6 Astra, Claude Opus 5, and Opus 5.5 across five reasoning levels on Lean 4 proof tasks and found diminishing returns kick in fast.

A planner-executor framework that coordinates parallel reasoning branches instead of independent sampling, lifting pass@64 on math benchmarks by up to 13.4 points.

Ten Claude Opus 5.5 agents collaborated on a message board for 15 hours to produce C-HD, a Lean-verified shortest-path algorithm that beats published bounds in a specific density regime.

Microsoft Research shows removing the central orchestrator lets thousands of coding agents self-organize, lifting test-pass rates by up to 21 points.

A Microsoft Research and UC Berkeley team shows agents sharing a scratchpad beat parallel independent runs, setting new records on ARC-AGI-3, polyomino packing, and MNIST compression.

Kyutai's Voice of Reason turns GLM-4-Voice into a speech-native math solver, jumping GSM8K accuracy from 27.3% to 77.1% with no added latency.

Epoch AI marks the first Major Advance on its unsolved-math benchmark, with GPT-6 Astra driving the proof in an interactive session with three mathematicians.

Sakana AI launches the Frontier Intelligence Group, a research collective betting that Transformers and scaling laws aren't the final answer to AGI.

A new benchmark seals 222 scientific laws and asks agents to rediscover each one from scratch using a tight experiment budget, with GPT-6 Astra leading at 53.2%.

A new architecture proposes latent reasoning that grows with sequence length, sharing one recurrent state across prompt, response, training, and RL replay.

A math benchmark built to resist AI just fell. GPT-6 Astra cracked the final Tier 4 problem, closing out a 98 percent run in 14 months.

A new recurrent reasoning method mixes denoising with looped hidden states, hitting 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 with just 7M parameters.

Together AI highlights that Moonshot's open-weight Kimi K3 outscores Anthropic's newer Claude Fable 5.1 by 60% on the hard slice of Harvey's autonomous legal agent benchmark.

A new evolutionary model shows that when computation, replication, and social behavior all draw from one energy budget, cooperation emerges naturally in populations of random Z80 programs.

A new paper shows recurrent-depth reasoning models behave like chaotic dynamical systems, where hard problems create fractal basins that trap thinking near wrong answers.

An internal OpenAI model coordinated roughly 10,000 agents for 88 hours to produce a Lean-verified finite-time blowup proof for 3D Navier-Stokes.

A new physics inspired bound explains which patterns stochastic gradient descent learns first, tying acquisition speed to Fisher information flow.

Meta's next-generation autonomous research agent AIRA₃ placed 8th of roughly 4,000 human teams in an NVIDIA-run Kaggle contest to fine-tune Nemotron.
No stories match these filters.