
PrismML Squeezes Qwen3.8 27B Into 5.9 GB With 98% Performance Retained
PrismML's ternary-quantized 27B model retains 98.2% of full-precision Qwen3.8 27B performance in a 5.9GB footprint, hitting 143 tokens/sec on an RTX 5090.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

PrismML's ternary-quantized 27B model retains 98.2% of full-precision Qwen3.8 27B performance in a 5.9GB footprint, hitting 143 tokens/sec on an RTX 5090.

Anthropic published three transparency metrics tracking AI-driven R&D, agent oversight, and compute allocation, urging other frontier labs to adopt the same reporting.

OpenAI unveils a legal-specific configuration of GPT-6 Astra with a 230M-URL search index, firm-built workflows, and 73 plugins for practitioners.

Krea's creative agent gains a timeline editor that lets you stitch generated clips, extend shots, and layer audio without leaving the session.

Vals AI released Vibe Code Bench 1-100, a benchmark that measures whether coding agents can extend a working web app across up to ten sequential product requests without breaking it.

Figure's new humanoid policy walked into 30 unseen Bay Area homes and tidied, folded, and made beds with zero prior data from those spaces.

Google Labs is turning CC from a solo daily-briefing agent into a shared household coordinator with its own Google account and permissions.

Anthropic opens vetted access to its restricted Mythos model for biology work, swapping real-time blocking for offline monitoring under a shared-responsibility framework.

Pika relaunches as a multi-app creative platform, bundling top video and image models with auto-selection and aggressive pricing on Seedance generations.

Exa Snapshot pins search and page retrieval to any past date, using 400 billion stored webpages to prevent evaluation leakage and enable backtests.

Claude Code projects now run parallel cloud threads coordinated by a chief-of-staff style controller, with shared memory that persists across sessions.

Goodfire's activation probes catch AI models cheating in real time, cutting monitoring costs 90% while flagging hacks that chain-of-thought judges miss.

Warp Factories now ships Scorers, LLM-as-judge agents that grade coding agent runs on quality, efficiency, and compliance to drive automatic self-improvement.

Runway's new Enhance Frame Rate model retimes video to any target between 24 and 120 fps, including broadcast fractional rates, priced by input seconds.

Perplexity Computer now lets users pick effort levels from Light to Ultra, and the orchestrator picks the right model and reasoning depth automatically.

Jina AI released jina-ocr-v1, a 3.4B mixture-of-experts document parser with speculative decoding that turns pages into Markdown at 2.57 pages per second.

vLLM's latest optimizations push Kimi K3 serving to 2.2 to 2.8x higher throughput on B300 GPUs, with TTFT cut by up to 85%.

Liquid AI and Insilico Medicine released two small LFM2 variants that beat GPT-5, Gemini-3.1-Pro, and Claude Opus on aging biology benchmarks.

Zed's weekly release adds animated cursors, Emmet wrap-with-abbreviation, configurable window titles, and Markdown files that open straight into rendered preview.

A community imatrix-quantized GGUF build of an uncensored GLM-4.7-Flash fine-tune brings local Chinese and English roleplay to consumer hardware.

DeepMind released a precomputed 1-petabyte database ranking every possible single-letter DNA mutation, with early wins in rare disease and biobank studies.

Browser Use released Jev Ultrafast, a tiny open source browser agent that books a Google Flights search in 7 seconds for under half a cent.

Zenbu Labs shipped a Chromium-powered browser that renders inside your terminal, giving coding agents a real web surface they can drive alongside your code.

Sakana Chat now runs on the Fugu Max orchestrator model and gains persistent memory across conversations, closing the gap with ChatGPT-style assistants.
No stories match these filters.