
Anthropic's Claude Rewrites 36 Biology AI Tools to Run 4x Faster
Claude accelerated over 30 open-source biology models by roughly 4x, letting a single GPU node predict structures once out of reach.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Claude accelerated over 30 open-source biology models by roughly 4x, letting a single GPU node predict structures once out of reach.

PrismML's ternary-quantized 27B model retains 98.2% of full-precision Qwen3.8 27B performance in a 5.9GB footprint, hitting 143 tokens/sec on an RTX 5090.

Prism ML's Bonsai 2 27B compresses a 27B reasoning model to 8.6 GB using ternary weights, retaining 98.2% of FP16 benchmark quality.

Empero distilled Qwen3.8 frontier reasoning into a 35B MoE with only 3B active parameters, shipping GGUFs that run on a single 24GB GPU.

Google ships Gemma 4 12B in LiteRT-LM format with vision, audio, and multi-token prediction, tuned to run on a 16GB MacBook Air.

Perplexity's local agent stack now runs on Windows RTX PCs, with on-device MCP servers and scheduled task automation joining the release.

Edge0 releases an 8B sparse MoE that runs in under 1 GiB of active memory on Apple Silicon, streaming experts from SSD on demand.

A community quantization squeezes DeepSeek's 552B multimodal MoE onto two desktop DGX Sparks, holding decode speeds across half a million tokens of context.

Edge0 streams a 35B Mixture-of-Experts model from SSD on an iPhone, holding under 3 GB of active RAM while decoding at 15 tokens per second.

Alibaba's 27B dense multimodal model lands on Cerebras at roughly 1,800 tokens per second, with reasoning on by default and a 128K context on paid tiers.
PyTorch 2.14 lands with a new CUTLASS-based GEMM backend, in-place fault tolerance for distributed training, and native linear algebra on Apple Silicon.

Together AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.

Epoch AI's new explorer estimates compute capacity for OpenAI, Google DeepMind, Anthropic, Meta, and xAI, revealing a 17x surge at OpenAI.

Cognition used Devin agents to build a GPU lattice siever that factored RSA-260, cutting factorization costs by roughly 10x versus prior public state of the art.

An autonomous agent ran 111 trials to optimize vLLM for Qwen3.8-27B on a single RTX 5090, hitting 4.1x throughput at 65K context.

NVIDIA ships CUDA Python 1.0 with semantic versioning, a shared cuda.core foundation, and PyTorch and CuPy already building on it.

An open-source patch restores 64GB of HBM2e memory and full compute on NVIDIA's crippled CMP 170HX mining card, transforming a $200 crypto relic into a viable AI accelerator.

OpenBMB's compact 2B model hits open-source SOTA in its class, beating several 4B rivals on code, math, and agent benchmarks while running locally.

A unified pruning, quantization and distillation pipeline shrinks a Vision Transformer 54.5x while holding 95.13% accuracy on out-of-distribution chilli disease images.

A new attention kernel unlocks Blackwell's FP4 tensor cores for inference, hitting 2.13x BF16 forward throughput on GB200 while training stays partly in FP8.

Figure locks in a $3.5B compute commitment scaling past $6B, targeting up to 100,000 Vera Rubin GPUs to train its Helix humanoid AI.
The latest PyTorch release ships NVGEMM CUTLASS kernels for Inductor, a rebuilt nccl2 backend, native Apple Silicon linear algebra, and first-class fault tolerance.

A community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.

A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.
No stories match these filters.