
Runway's Ruby Rebuilds SDR Footage Into True HDR for Pro Workflows
Runway's new Ruby model upconverts standard dynamic range footage to true 16-bit HDR in ProRes and EXR, targeting professional post-production pipelines.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Runway's new Ruby model upconverts standard dynamic range footage to true 16-bit HDR in ProRes and EXR, targeting professional post-production pipelines.

Artificial Analysis launched a blind human-preference leaderboard for voice agents, and the model users like most is not the one that finishes the task.

An open-source tool parses your monorepo with Tree-sitter, builds a Memgraph knowledge graph, and hooks into Claude Code via MCP.

TrueFoundry open-sourced TrueForge, an MIT-licensed agent runtime that ran 30-75% cheaper than Claude Managed Agents on a 14-task enterprise benchmark.

NVIDIA's AVO agent architecture lifts Claude Opus 5 from a 30% baseline to a perfect 100 on ARC-AGI-3's 183 interactive reasoning levels.

DeepSeek open-sourced dsh, a plugin-based coding agent framework where the model, tools, sandbox, UI, and even the agent loop itself are all swappable.

A new agent skill turns a single reference image into diffable TypeScript that procedurally reconstructs the object as an animation-ready Three.js scene.

Block open-sourced Berd, a Tauri-based desktop shell that unifies AI agents, projects, and files across any model provider under Apache 2.0.

A single-file C engine streams experts from disk to run 744B-parameter MoE models on a 25GB laptop, no GPU required.

An unofficial Tauri 2 desktop app wraps xAI's Grok Build CLI with sessions, project management, media previews, and scheduled automations.

Hugging Face released Tau, an MIT-licensed terminal coding agent designed as much for reading as running, with a three-layer architecture you can actually understand.

A Go-based open source relay platform pools Claude, OpenAI, Gemini, and Grok subscriptions behind one endpoint with billing, sticky sessions, and cost sharing.

DeepSeek's new experimental vision model bolts image understanding onto V4-Flash at the same price, closing in on Claude Opus 4.8 on multimodal agent tasks.

Stanford's Marin lab kicked off a 535B mixture-of-experts run with 23B active parameters, streaming every metric, config, and mistake live to the public.

A Swift CLI from LY Corporation gives AI coding agents a token-efficient observe-act loop across iOS Simulator and Android devices.

Sokuji is an open-source live speech translator that runs entirely on-device with WASM and WebGPU, no API key or GPU required.

Magnitude's new catalog profiles your hardware, estimates tokens per second for every model, and picks the best local LLM before you download anything.

Sakana AI upgraded its free JP-EN-ZH translator to the new Namazu model, beating Google Translate, DeepL, and Claude Opus 4.8 in head-to-head evaluations.

A community-built vLLM container brings NVFP4 KV cache, DFlash speculative decoding, and Blackwell sm_121a runtime patches to DGX Spark serving.

Alibaba's third-generation image model debuts at #6 in editing and #9 in text-to-image, with big Elo jumps and a productivity-first pitch.

Anthropic makes computer use, browser tool, Skills API, and Files API generally available with batched actions cutting round trips 20-40%.

GPT-Image-2 can now generate PNGs with real alpha channels directly through the API, skipping the separate background-removal step for cutouts.

Kaggle and Gert Labs turned identity-theft prevention into a two-model roleplay, testing whether LLMs can catch social engineers without stonewalling real customers.

A per-tensor dynamic quantization of an abliterated Qwen3.8-27B lands on Hugging Face, delivering 82.98% MMLU and 262k context on a single RTX 4090.
No stories match these filters.