
Google DeepMind's WeatherNext 3 Ditches Physics Simulations for 60% Sharper Forecasts
Google DeepMind's new weather AI trains on raw satellite feeds and station data, delivering hourly 5km forecasts with up to 50% better rain accuracy.
A focused feed of models, agents, research and open-source releases for people building with AI.

Google DeepMind's new weather AI trains on raw satellite feeds and station data, delivering hourly 5km forecasts with up to 50% better rain accuracy.

Xiaohongshu's Self-GC uses a planner LLM to fold, mask, or prune agent context, retaining critical details 84.85% of the time versus 54.55% for rule-based baselines.

An 8-year interpretability project shows that LLM representations can be closely approximated by symbolic role-filler structures, enabling precise behavioral edits.

Google's third Flash release in as many months pushes a workhorse model into frontier-tier territory on agentic coding, legal, and finance benchmarks.

Multiverse Computing's Quasar 438B tops European AI rankings with a 43 Intelligence Index score and 15.3-second reasoning responses.

Alibaba's flagship 2.4T-parameter model gets a coding and agent-work refresh with the same pricing and 1M context window.

Independent evaluations put Anthropic's new flagship at the top of the intelligence charts, but the win comes with a token-count tax that eats the cache savings.

A new benchmark study finds that sliding-window attention with sinks matches or beats post-trained linear attention models, without any retraining.

Cursor added Anthropic's new Fable 5.1 model, which posted a 73.4% score on CursorBench 3.2 and excels at self-verifying long coding runs.

Shieldstral accepts policies at runtime. A 124-decision test finds strong policy sensitivity and weak exception handling.

New scaling laws show that looping the middle half of a Mixture-of-Experts model twice saves up to 18% of training compute at matched budgets.

A community fine-tune of Qwen3.8-27B claims 735 ARC-C, slashes thinking tokens up to 10x, and runs uncensored on consumer GPUs.

Google's new 330M-parameter time series foundation model handles multivariate forecasting zero-shot in a single forward pass, topping Gift-Eval, FEV-Bench and Time.

Seven practical shifts for securing agents that can find any crack

A double-refined abliteration of Qwen3.8-27B cuts refusals to near zero while shrinking behavioral damage roughly sixfold versus its upstream.

A community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.

DeepSeek gives its Flash model eyes with an experimental multimodal release that matches Claude Opus 4.8 on agent tasks at a fraction of the cost.

Z.ai shipped a frontier coding model without touching the base weights, matching GPT-5.6 Sol and Claude Fable 5 on agentic tasks through scaled post-training alone.

Apodex 1.1 lands with an Elo of 1348 on GDPval-AA v2, beating DeepSeek V4 Pro and Kimi K2.6 on real-world agentic work.

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

OrcaRouter stripped the refusal alignment from Z.ai's 320B GLM-5.3-Flash MoE, baking the edit directly into the official block-FP8 shards.

An AI assistant now handles online checkout end to end by connecting to Stripe Link and using single-use virtual cards that need per-purchase approval.

Perplexity Computer now routes long-context, multimodal research tasks to GLM 5.3, which outperformed GLM 5.2 on the in-house WANDR benchmark.

Epoch AI's EBR-bench human baseline shows people quickly outclass frontier models at Earthborne Rangers, exposing a real learning gap.
No stories match these filters.