
Alibaba's Qwen3.8-Omni-Flash Cuts Video AI Costs by 89% With Agent Tool Use
Alibaba's new omni-modal model reasons over audio and video, orchestrates tools across long workflows, and cuts video input costs by roughly 89%.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Alibaba's new omni-modal model reasons over audio and video, orchestrates tools across long workflows, and cuts video input costs by roughly 89%.

Pika relaunches as a multi-app creative platform, bundling top video and image models with auto-selection and aggressive pricing on Seedance generations.

Runway's new Enhance Frame Rate model retimes video to any target between 24 and 120 fps, including broadcast fractional rates, priced by input seconds.

Runway's Fall 2026 drop bundles Fish Audio S2.1 Pro, MiniMax H3 Max, Cartesia Sonic 3.6, plus Flux video upscaling and editing into one workspace.

Odyssey unveils a single foundation world model that controls robots, humanoids, cars, drones, and video game characters with only hours of task data.

Hypit gives coding agents like Claude Code and Codex a video language that binds shots, captions, and B-roll to spoken words instead of timeline seconds.

A viral open-source toolkit turns any novel into character bibles, outlines, art references, scripts, and storyboards using Claude Code or Codex.

NVIDIA and USC present HorizonRelight, a diffusion transformer method that keeps lighting stable across long videos by propagating context between chunks.

Runway details how it turns diffusion video models into causal, autoregressive streamers using teacher forcing and on-policy distillation.

Black Forest Labs shipped a prompt-driven video editor that changes one thing per clip while leaving length, framing, and audio untouched.

CozyClay turns your browser into a previs studio where you block shots, pose characters, and hand the same camera moves to an AI video model.

A new open-source Claude Code and Codex skill turns any topic into a fully coded motion-graphics explainer video with voiceover, subtitles and QC.

FastVideo's 8-step distilled MiniMax H3 checkpoint lands as a single-file ComfyUI repack, bringing text-to-video-with-audio generation to local workflows.

World Labs unveiled Atlas, an omni world model that unifies text, images, video, and 3D into a single 3D-grounded spatial context.

Krea's new Realtime Director lets you steer a live AI video stream with prompts as it plays, powered by fal's H3 Max Director model.

Unitree's UnifoLM-X2-1.0 lets humanoid robots spar autonomously by predicting the next few seconds of physics in real time.

VDN-H3 replaces most of MiniMax H3's quadratic attention with a linear branch, rendering a 14.4-second 768p clip in 11.23 seconds on 8 B200s.

A community LoRA for MiniMax H3 sharpens video output through aligned guide latents, offering a new pass on ComfyUI's ref2va pipeline.

Runway's new general world model streams continuous 720p video and 48kHz audio in real time, taking text prompts and camera input as you play.

A ComfyUI custom node makes MiniMax H3 clips chain seamlessly by carrying motion and the exact audio waveform across cuts.

Amap CV Lab open-sourced ABot-Recon, a streaming 3D reconstruction model that runs on a rolling 12-frame window at 24 FPS.

A new open-source agent skill hands Claude Code and Codex a full motion-design studio for voiceover explainer videos with word-level sync and Remotion rendering.

Utopia is a self-hosted Rust and Postgres knowledge platform that layers RAG on top of a bitemporal knowledge graph tracking when each fact was true.

Runway Ruby now exports scene-referred half-float EXR sequences in ACEScg 1.3 and 2.0, slotting AI video directly into professional color pipelines.
No stories match these filters.