
Prism ML's Ternary Bonsai 2 Squeezes a 27B Reasoning Model Into 8.6 GB
Prism ML's Bonsai 2 27B compresses a 27B reasoning model to 8.6 GB using ternary weights, retaining 98.2% of FP16 benchmark quality.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Prism ML's Bonsai 2 27B compresses a 27B reasoning model to 8.6 GB using ternary weights, retaining 98.2% of FP16 benchmark quality.

A fresh 66-task suite pushes agents beyond software into hardware, science, and media, with only three models clearing 30 percent.

OpenAI publishes a voluntary disclosure process for misalignment findings and drops six case studies of models cheating, hiding mistakes, and coordinating without permission.

Grok Bot now plugs into 1Password so its cloud browser can log into any site without your agent ever seeing the raw password.

A new open-source launcher routes Codex through your ChatGPT Web session, letting Pro subscribers use their chat quota for agentic coding instead of API credits.

A team from Google, Google DeepMind, University of Maryland, and University of Virginia turns past search trees into cheap replay simulators for meta-exploration.

Empero distilled Qwen3.8 frontier reasoning into a 35B MoE with only 3B active parameters, shipping GGUFs that run on a single 24GB GPU.

A new benchmark seals 222 scientific laws and asks agents to rediscover each one from scratch using a tight experiment budget, with GPT-6 Astra leading at 53.2%.

Google ships Gemma 4 12B in LiteRT-LM format with vision, audio, and multi-token prediction, tuned to run on a 16GB MacBook Air.

Baseten's research arm teams up with Hugging Face and Goodfire to embed safety controls into open-weight models from training through runtime serving.

NVIDIA's Axolotl3D fuses images, camera poses, and partial point clouds into one diffusion pipeline that completes occluded 3D shapes faithfully.

Epoch AI's data center explorer now maps 86 sites covering an estimated 44% of global AI compute, with satellite imagery and detailed hardware specs.

Runway's Fall 2026 drop bundles Fish Audio S2.1 Pro, MiniMax H3 Max, Cartesia Sonic 3.6, plus Flux video upscaling and editing into one workspace.

Anthropic is folding its agentic workspace into the main Claude app and launching Docs and Slides in beta, so any chat can now spin off long-running work.

Cohere is adding hardware-backed confidential computing to Model Vault, encrypting prompts and model activations inside NVIDIA GPUs during inference itself.

Cognition extended Devin with Code Scans, codebase-wide audits powered by Agentic MapReduce that investigate, report findings, and open pull requests automatically.

The San Francisco lab that trained a 400B open-weight model for $20M just raised $150M at a unicorn valuation to chase China.

Toronto's Cohere and Heidelberg's Aleph Alpha finalized their merger, forming a $20B transatlantic company aimed at sovereign AI for governments and regulated industries.

Zed's new Delta environment replaces GitHub pull requests with agent-aware threads, backed by a version control system that records every edit.

Ant Group's Ling-3.0-flash-Fin is a 124B MoE reasoning model tuned for financial research, released open weights under MIT with a 256K context.

A preregistered experiment with 240 people shows that chatting with an LLM trained only on pre-1930 text erases the illusion that the past was more moral.

Mistral's models will power Firefox Smart Window, Mozilla's opt-in AI browsing mode, starting in France and North America with zero data retention.

China Telecom released Xing4.0-29B-A4B , a 29B MoE with 4B active parameters, Apache 2.0. Native 256K context (extensible to 512K) using MLA attention plus multi-token prediction heads. First model of this scale trained entirely on Huawei Ascend 910C NPUs with MindSpore.

A community developer stripped refusal behavior from Qwen3.8-27B using Heretic's automated abliteration, keeping benchmarks within noise while cutting refusals from 98/100 to 12/100.
No stories match these filters.