
Perplexity's Q2D-Web Benchmark Tests AI Search on 190M Real Web Documents
Perplexity released Q2D-Web, a large-scale benchmark with 190M documents and 70K agent-reformulated queries for evaluating retrieval in agentic RAG systems.
A focused feed of models, agents, research and open-source releases for people building with AI.

Perplexity released Q2D-Web, a large-scale benchmark with 190M documents and 70K agent-reformulated queries for evaluating retrieval in agentic RAG systems.

Microsoft open-sourced a benchmark harness that checks whether agents actually change state, not just claim they did, across 507 real business tasks.

A new paper shows recurrent-depth reasoning models behave like chaotic dynamical systems, where hard problems create fractal basins that trap thinking near wrong answers.

OpenBMB's 2.6B dense reasoning model tops the sub-4B open weights leaderboard, punching well above its weight on agentic benchmarks while staying token-efficient.

NVIDIA's fine-tuned Nemotron system scored 535.4 out of 600 at IOI 2026, topping the best human contestant under identical contest conditions.

Artificial Analysis pushes Intelligence Index v4.2 with private test sets, a 4,592-page PDF reasoning benchmark, and a new agentic knowledge-work suite.

LlamaIndex and Kaggle launch a schema-guided document extraction benchmark that grades models on missing fields, source grounding, and repeated records across 370 enterprise files.

Claude autonomously wrote a 13 million line Lean proof of Fermat's Last Theorem in 11 days, verifying 29,500 supporting theorems along the way.

GitHub's new research preview routes each coding task through single, cascade, or critique workflows across multiple models, beating Opus 5 at a fraction of the cost.

Warp launches Factory Benchmarks, letting teams replay real agent runs to pick the best model per task and cut cost-per-PR from $80 to $30.

Microsoft AI's new speech recognition model hits 2.0% word error rate at 411x real time speed for $1.67 per 1,000 minutes.

A new generative retriever leverages the table of contents to hit 82.6% Recall@1 on a fresh 18-book benchmark, with hallucinations under 0.05%.

A new 365-day simulation drops LLM agents into a deterministic marketplace with ¥100,000, real supplier data, and fraudsters, and no model wins everything.

Artificial Analysis rebuilt its Image Editing Arena with a two-axis taxonomy of editing actions and use cases, revealing which model wins each specific task.

Multiverse Computing's Quasar 438B tops European AI rankings with a 43 Intelligence Index score and 15.3-second reasoning responses.

Independent evaluations put Anthropic's new flagship at the top of the intelligence charts, but the win comes with a token-count tax that eats the cache savings.

Lovable's new default coding model gains 17% on the hardest fix-and-iterate tasks while cutting cost by 31%, with sharper self-verification and visual taste.

Meta Superintelligence Labs' new real-time speech model tops streaming ASR and diarization benchmarks, handles 20+ speakers, and ships at $0.18 per hour.

Google's new 330M-parameter time series foundation model handles multivariate forecasting zero-shot in a single forward pass, topping Gift-Eval, FEV-Bench and Time.

Z.ai shipped a frontier coding model without touching the base weights, matching GPT-5.6 Sol and Claude Fable 5 on agentic tasks through scaled post-training alone.

Apodex 1.1 lands with an Elo of 1348 on GDPval-AA v2, beating DeepSeek V4 Pro and Kimi K2.6 on real-world agentic work.

Broad Institute researchers unveil a framework that separates score-chasing from real scientific reasoning, exposing where Claude, GPT, and Gemini break down.

Perplexity's new Search API takes the top three spots on the Artificial Analysis Search Index, extending the quality-cost Pareto frontier for agentic search.

Epoch AI's EBR-bench human baseline shows people quickly outclass frontier models at Earthborne Rangers, exposing a real learning gap.
No stories match these filters.