
Claude Prompt Caching: 6 Tricks That Cut Your API Bill
Breakpoints, TTL math, Batch stacking, and a two-session split. Measured on OpenRouter, $0.003 vs $0.022 per turn
A focused feed of models, agents, research and open-source releases for people building with AI.

Breakpoints, TTL math, Batch stacking, and a two-session split. Measured on OpenRouter, $0.003 vs $0.022 per turn

An autonomous agent ran 111 trials to optimize vLLM for Qwen3.8-27B on a single RTX 5090, hitting 4.1x throughput at 65K context.

NVIDIA ships CUDA Python 1.0 with semantic versioning, a shared cuda.core foundation, and PyTorch and CuPy already building on it.

Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.

vLLM's new Hybrid HiSparse keeps long-context requests decoding when the KV cache overflows HBM, tripling concurrency on GLM 5.3 at 1M context.

Experiential is an open source, OpenAI-compatible gateway that routes across hosted, BYOK, and local models with zero markup and spend controls.

Perplexity open-sources details of Ivy, Tulip, and ROSE, its Rust plus Python embedding stack that beats vLLM on latency and throughput.

Prime Intellect rebuilt weight sync on RDMA and vLLM tracing, dropping GLM-5.2's 1.6TB policy transfer from 86 seconds to under 4.

Hermes Desktop now auto-detects your hardware, picks a fitting local model, downloads it, and configures the runtime without any manual setup.

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

Google DeepMind's new weather AI trains on raw satellite feeds and station data, delivering hourly 5km forecasts with up to 50% better rain accuracy.

Cursor now lets teams run cloud coding agents on their own infrastructure or partnered sandbox providers, with auto-scaling worker pools.
Perplexity open-sources Lily, a Metal-based inference engine tuned for Qwen3.6-35B-A3B that beats MLX-LM by 1.23x prefill and 1.35x decode on M5 Max.

A new benchmark study finds that sliding-window attention with sinks matches or beats post-trained linear attention models, without any retraining.

Perplexity's Computer now splits agent tasks between cloud frontier models and an on-device model, gated by an open-source 0.6B PII detector.

Kimi Code 0.39.0 ships an experimental Remote Control mode that lets you drive a local coding session from any browser or phone.

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.

Red Hat AI shipped an NVFP4 quantization of Z.ai's 320B GLM-5.3-Flash, shrinking the model to run on vLLM with FP4 activations while holding reasoning benchmarks near the original.

The latest vLLM release lands 584 commits from 270 contributors, with big performance wins for Kimi-K3, DeepSeek-V4, and speculative decoding.

A new method throws away most of a model's own reasoning trace mid-thought, cutting memory to a fixed cap and running inference 3x faster.

Cerebras unveiled the CS-4 wafer-scale system architecture at Hot Chips and sketched a roadmap to CS-5 and 3D-stacked DRAM in CS-6.

NVIDIA Dynamo's new shadow engine keeps a warm standby ready on the same GPU, cutting LLM failover from minutes to seconds.

Apple officially endorsed exo on its Mac Studio and Mac Mini pages, blessing a workflow that turns four desktops into a 4.8TB/s inference rig.

OpenAI's first in-house inference chip delivers 1.5-1.9x more work per watt and up to 4.1x lower latency versus current systems.
No stories match these filters.