
Cursor Lets Enterprise Teams Run AI Coding Agents on Their own Servers
Cursor now lets teams run cloud coding agents on their own infrastructure or partnered sandbox providers, with auto-scaling worker pools.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Cursor now lets teams run cloud coding agents on their own infrastructure or partnered sandbox providers, with auto-scaling worker pools.
Perplexity open-sources Lily, a Metal-based inference engine tuned for Qwen3.6-35B-A3B that beats MLX-LM by 1.23x prefill and 1.35x decode on M5 Max.

A new benchmark study finds that sliding-window attention with sinks matches or beats post-trained linear attention models, without any retraining.

Perplexity's Computer now splits agent tasks between cloud frontier models and an on-device model, gated by an open-source 0.6B PII detector.

Kimi Code 0.39.0 ships an experimental Remote Control mode that lets you drive a local coding session from any browser or phone.

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.

Red Hat AI shipped an NVFP4 quantization of Z.ai's 320B GLM-5.3-Flash, shrinking the model to run on vLLM with FP4 activations while holding reasoning benchmarks near the original.

The latest vLLM release lands 584 commits from 270 contributors, with big performance wins for Kimi-K3, DeepSeek-V4, and speculative decoding.

A new method throws away most of a model's own reasoning trace mid-thought, cutting memory to a fixed cap and running inference 3x faster.

Cerebras unveiled the CS-4 wafer-scale system architecture at Hot Chips and sketched a roadmap to CS-5 and 3D-stacked DRAM in CS-6.

NVIDIA Dynamo's new shadow engine keeps a warm standby ready on the same GPU, cutting LLM failover from minutes to seconds.

Apple officially endorsed exo on its Mac Studio and Mac Mini pages, blessing a workflow that turns four desktops into a 4.8TB/s inference rig.

OpenAI's first in-house inference chip delivers 1.5-1.9x more work per watt and up to 4.1x lower latency versus current systems.

Perplexity's new Portable Computer runs the full agent stack locally on NVIDIA DGX Spark, with cloud escalation gated by user approval and zero per-token cost for on-device work.

OpenAI now lets teams attribute spend down to individual API keys and enforce hard monthly caps that cut off traffic when hit.

Pika Speech is a 3B flow-matching TTS model that generates a minute of 48 kHz audio in about a second at one cent per minute.

A single-file C engine streams experts from disk to run 744B-parameter MoE models on a 25GB laptop, no GPU required.

A Go-based open source relay platform pools Claude, OpenAI, Gemini, and Grok subscriptions behind one endpoint with billing, sticky sessions, and cost sharing.

A community-built vLLM container brings NVFP4 KV cache, DFlash speculative decoding, and Blackwell sm_121a runtime patches to DGX Spark serving.

Liquid AI ships DSpark draft models for its LFM2.5 family, delivering up to 3.18x GPU throughput and cutting agent latency by nearly half.
No stories match these filters.