
When to Gate Recursive Terminal-task Synthesis
Mark each agent job Promote, Hybrid, or Reject. RST is allowed only when the job is hermetic, bounded, reproducible, and outcome-verifiable
A focused feed of models, agents, research and open-source releases for people building with AI.

Mark each agent job Promote, Hybrid, or Reject. RST is allowed only when the job is hermetic, bounded, reproducible, and outcome-verifiable

A new physics inspired bound explains which patterns stochastic gradient descent learns first, tying acquisition speed to Fisher information flow.

Prime Intellect rebuilt weight sync on RDMA and vLLM tracing, dropping GLM-5.2's 1.6TB policy transfer from 86 seconds to under 4.
The latest PyTorch release ships NVGEMM CUTLASS kernels for Inductor, a rebuilt nccl2 backend, native Apple Silicon linear algebra, and first-class fault tolerance.

A new frequency-domain training objective for pixel-space flow matching cuts convergence time by up to 40% without touching the architecture.

A Tsinghua team trained a 2B language model from scratch on consumer RTX 5090s for under $6,900, matching Qwen2.5-1.5B and open-sourcing everything.

New scaling laws show that looping the middle half of a Mixture-of-Experts model twice saves up to 18% of training compute at matched budgets.

A one-line change to LoRA's initialization closes the gradient gap with full fine-tuning, boosting speed and stability with zero extra parameters.

Brett Adcock's stealth AI startup lands gigawatt-scale Vera Rubin capacity, joining Thinking Machines and xAI in NVIDIA's biggest compute alliances.

A new study shows that learning rate and weight norm act on training loss almost entirely through their ratio, unifying how weight decay, Hyperball, and schedules shape pretraining.

SkyRL's IsoExec makes vLLM rollout and Megatron training agree bitwise on token log-probabilities, cutting numerical drift below 1e-6 with 25% overhead.

Stanford's Marin lab kicked off a 535B mixture-of-experts run with 23B active parameters, streaming every metric, config, and mistake live to the public.

Kakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.
No stories match these filters.