
Perplexity Brings Portable Computer to Windows RTX PCs for Free Local AI
Perplexity's local agent stack now runs on Windows RTX PCs, with on-device MCP servers and scheduled task automation joining the release.
A focused feed of models, agents, research and open-source releases for people building with AI.

Perplexity's local agent stack now runs on Windows RTX PCs, with on-device MCP servers and scheduled task automation joining the release.

Edge0 releases an 8B sparse MoE that runs in under 1 GiB of active memory on Apple Silicon, streaming experts from SSD on demand.

A community quantization squeezes DeepSeek's 552B multimodal MoE onto two desktop DGX Sparks, holding decode speeds across half a million tokens of context.

Edge0 streams a 35B Mixture-of-Experts model from SSD on an iPhone, holding under 3 GB of active RAM while decoding at 15 tokens per second.

Alibaba's 27B dense multimodal model lands on Cerebras at roughly 1,800 tokens per second, with reasoning on by default and a 128K context on paid tiers.
PyTorch 2.14 lands with a new CUTLASS-based GEMM backend, in-place fault tolerance for distributed training, and native linear algebra on Apple Silicon.

Together AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.

Epoch AI's new explorer estimates compute capacity for OpenAI, Google DeepMind, Anthropic, Meta, and xAI, revealing a 17x surge at OpenAI.

Cognition used Devin agents to build a GPU lattice siever that factored RSA-260, cutting factorization costs by roughly 10x versus prior public state of the art.

An autonomous agent ran 111 trials to optimize vLLM for Qwen3.8-27B on a single RTX 5090, hitting 4.1x throughput at 65K context.

NVIDIA ships CUDA Python 1.0 with semantic versioning, a shared cuda.core foundation, and PyTorch and CuPy already building on it.

An open-source patch restores 64GB of HBM2e memory and full compute on NVIDIA's crippled CMP 170HX mining card, transforming a $200 crypto relic into a viable AI accelerator.

OpenBMB's compact 2B model hits open-source SOTA in its class, beating several 4B rivals on code, math, and agent benchmarks while running locally.

A unified pruning, quantization and distillation pipeline shrinks a Vision Transformer 54.5x while holding 95.13% accuracy on out-of-distribution chilli disease images.

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

Figure locks in a $3.5B compute commitment scaling past $6B, targeting up to 100,000 Vera Rubin GPUs to train its Helix humanoid AI.
The latest PyTorch release ships NVGEMM CUTLASS kernels for Inductor, a rebuilt nccl2 backend, native Apple Silicon linear algebra, and first-class fault tolerance.

A community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.

A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.

A 26B multimodal Gemma variant with refusal directions surgically removed lands on Hugging Face in GGUF format, ready for llama.cpp.

A community fine-tune of Qwen3.8-27B strips refusals, fixes the resulting weight damage, and ships in four formats tuned for local inference.

A community project squeezes Qwen3.8-27B onto a 24GB gaming card with vLLM, hitting 417 tok/s batched or 82 tok/s single-user at 150k context.
No stories match these filters.