
Comfy-Org Brings YuE2 Music Generation Natively Into ComfyUI Workflows
Comfy-Org has repackaged the MAP team's YuE2 music model into single-file safetensors, and ComfyUI now ships native nodes to run it.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Comfy-Org has repackaged the MAP team's YuE2 music model into single-file safetensors, and ComfyUI now ships native nodes to run it.

Sakana AI's Fugu Max and Fugu Ultra v2 orchestrate pools of open models to beat frontier LLMs on cost and capability simultaneously.

Google Research introduces ToolGrad, an answer-first framework that generates tool-use training data with a 99.8% pass rate at lower cost.

Cursor's new Projects feature replaces one-off chats with a persistent coordinator agent that delegates to subagents, syncs context, and runs recurring work.

NVIDIA and USC present HorizonRelight, a diffusion transformer method that keeps lighting stable across long videos by propagating context between chunks.

OpenAI opens up the same agent harness that powers Codex, letting developers spin up long-running cloud agents with a single API call.

Anthropic ships a live session viewer and a server-side auto permission mode for Claude Managed Agents, letting the model decide when to pause for human approval.

OpenAI launches a vertical ChatGPT for banks that bundles Daloopa, PitchBook, and LSEG data with GPT-6 Astra reasoning and firm-specific templates.

humans& releases Persimmon, a 550B user model built to simulate real people in group chats, fooling AI judges 20 percent of the time.

Runway details how it turns diffusion video models into causal, autoregressive streamers using teacher forcing and on-policy distillation.

Together AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.

A math benchmark built to resist AI just fell. GPT-6 Astra cracked the final Tier 4 problem, closing out a 98 percent run in 14 months.

OpenAI's full-duplex voice model lands in the API at $0.05 per minute, letting agents listen and speak simultaneously while delegating reasoning to backend models.

Anthropic's latest threat report details how state actors, criminals, and hacktivists weaponized Claude across cyber, surveillance, biology, and distillation campaigns.

A new recurrent reasoning method mixes denoising with looped hidden states, hitting 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 with just 7M parameters.

Edge0 is an open source framework that streams MoE experts from SSD, letting a 35B parameter model run on a 24GB Mac mini with under 3GB active memory.

Google folded the full Gemini API documentation directly into AI Studio, so developers can read reference material without leaving the build surface.

Cohere's new 218B MoE translation model tops WMT26 against DeepL, Google Translate, and open alternatives across 50+ languages under a non-commercial license.

Google launches a native Gemini desktop app for Windows 10 and 11 with an Alt+Space hotkey, Spark agent access, and image and video generation.

Google Pics packages Nano Banana into a Workspace-native image editor with object-level control, in-image text editing, and live collaboration.

Cognition's new coding model matches frontier systems while costing up to 70% less, driven by an RL recipe that trains every effort level at once.

Black Forest Labs shipped a prompt-driven video editor that changes one thing per clip while leaving length, framing, and audio untouched.

Krea's new agent platform reads your canvas, plans multi-model pipelines, and executes creative jobs from a single prompt, no node-wiring required.

Tencent Hunyuan opens the code and weights for a 1.5B unified speech model that handles TTS, editing, denoising, and separation from plain instructions.
No stories match these filters.