
Cohere Labs' North Small Translate Beats DeepL and Google Translate Across 50 Languages
Cohere's new 218B MoE translation model tops WMT26 against DeepL, Google Translate, and open alternatives across 50+ languages under a non-commercial license.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Cohere's new 218B MoE translation model tops WMT26 against DeepL, Google Translate, and open alternatives across 50+ languages under a non-commercial license.

Tencent Hunyuan opens the code and weights for a 1.5B unified speech model that handles TTS, editing, denoising, and separation from plain instructions.

Unitree open-sourced a 6B-parameter humanoid foundation model that runs 64 tasks with one network across grippers, dexterous hands, and full-body motion.

CozyClay turns your browser into a previs studio where you block shots, pose characters, and hand the same camera moves to an AI video model.

An open 3B music model from M-A-P matches Suno v5 on quality, adds editable ABC scores, and runs on a 24GB GPU.

Wake is a native macOS app that unifies coding-agent sessions from Claude Code, Codex, Cursor and nine more into one searchable window.

Cohere open-sourced a serving engine that runs the entire LLM decode step as one persistent CUDA kernel, hitting 1.58x vLLM throughput on H100.

An open-source Docker container runs Codex, Claude Code, and Hermes through one unified API, keeping your keys and data on your own machine.

vLLM's new Hybrid HiSparse keeps long-context requests decoding when the KV cache overflows HBM, tripling concurrency on GLM 5.3 at 1M context.

An open-source patch restores 64GB of HBM2e memory and full compute on NVIDIA's crippled CMP 170HX mining card, transforming a $200 crypto relic into a viable AI accelerator.

Adaption Labs plugs its automated fine-tuning system into Hugging Face, letting anyone push a co-optimized model to the Hub with a single click.

Nex-AGI releases an open-source agentic model family with vision-driven computer use, three sizes, and benchmark scores rivaling Claude Opus 5.

A hobbyist published a full DIY build kit for a Microduck-style RL biped robot, complete with a prebuilt SD image and MuJoCo training environment.

RadixArk's open-source Miles framework runs asynchronous RL on a 744B-parameter model across 64 GB300 GPUs with a 263-second median step.

Nex-AGI open-sourced a three-tier agentic model family under Apache 2.0, topped by a 1.6T-parameter MoE built to run browsers, desktops, and code loops end to end.

Red Hat's ripwire is a zero-dependency C++23 CLI that hands coding agents a ranked, deterministic map of any repo without embeddings or vector DBs.

AIPOCH released Open Science, an Apache-2.0 desktop workbench that runs scientific agents locally with Python notebooks, life-science connectors, and full provenance.

VDN-H3 replaces most of MiniMax H3's quadratic attention with a linear branch, rendering a 14.4-second 768p clip in 11.23 seconds on 8 B200s.

A community LoRA for MiniMax H3 sharpens video output through aligned guide latents, offering a new pass on ComfyUI's ref2va pipeline.

A 753B parameter GLM-5.3 variant has had its refusal circuits surgically removed for offensive security work, keeping MMLU intact while complying with 89% of cyber-attack prompts.

A Chinese developer's open-source Agent Skills portfolio bundles sixteen focused tools, including a layered Web Clipper that chains five extractors to reliably scrape anything.

Experiential is an open source, OpenAI-compatible gateway that routes across hosted, BYOK, and local models with zero markup and spend controls.

Alibaba's Zvec team open-sourced zg, a local-first CLI that fuses ripgrep, BM25, and vector search into one interface for humans and coding agents.
Perplexity open-sources Lily, a Metal-based inference engine tuned for Qwen3.6-35B-A3B that beats MLX-LM by 1.23x prefill and 1.35x decode on M5 Max.
No stories match these filters.