
World Labs' Atlas Beats Specialized 3D Models With One Omni Model
World Labs unveiled Atlas, an omni world model that generates 1440p video with pixel-perfect camera control and reconstructs scenes in 3D from a handful of images.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

World Labs unveiled Atlas, an omni world model that generates 1440p video with pixel-perfect camera control and reconstructs scenes in 3D from a handful of images.

JD.com open sources a 16B autoregressive diffusion editor that rewrites live video at 30 FPS in 720p from a text prompt.

Pireel is an open-source, browser-based AI video editor for talking-head content that any AI agent can drive over MCP.

Runway's Solaris skips the code layer entirely, generating live app interfaces frame by frame and beating top LLMs on visual reconstruction tests.

A new Gaussian splatting framework reconstructs volumetric videos with thousands of frames of complex motion, extending prior methods by roughly 70x in temporal coverage.

A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.

Sarvam opens its multilingual voice and video workspace to the public, bundling dubbing, cloning, and TTS across 11 Indian languages into one browser tool.

Google's video model gets scene extension to 40 seconds, keyframe control, 360p drafting, 4K upscaling, and video reference inputs.

Krea launches MiniMax H3 Max, a speed-tuned variant of the Hailuo 3 video model that renders 15 second clips in roughly 5 seconds.

A community developer shipped a 54-node ComfyUI extension for MiniMax H3 with same-day tooling for Alibaba PAI's new 8-step acceleration LoRAs.

FastVideo distilled MiniMax's 33B video-and-audio diffusion transformer from 50 denoising steps down to four, with open weights and 90% sparse attention.

A group of vision researchers argues that pure vision, not language-tethered multimodal models, could be its own route to general intelligence.

Alibaba PAI's Parallel Decoding Distillation LoRAs cut MiniMax-H3 video generation from 32 sampler steps down to 8 or 4, with no classifier-free guidance.

A style LoRA for MiniMax H3 generates video with hand-painted gouache backgrounds and golden-age character animation from text alone.

Pika launches a $10/month API Club offering 100+ generative media models through one endpoint, at up to 88% below rival aggregator prices.

Alibaba PAI drops a single ControlNet-Union checkpoint that adds Canny, Depth, HED, MLSD, Pose and inpainting to MiniMax-H3 video generation.

A community-made checkpoint grafts Z-Image's texture attention onto MiniMax H3's video engine, giving richer surfaces without retraining or extra VRAM.

Runway's new Ruby model upconverts standard dynamic range footage to true 16-bit HDR in ProRes and EXR, targeting professional post-production pipelines.

Black Forest Labs shipped a dedicated video super-resolution endpoint that regenerates any clip up to native 4K with two quality-vs-fidelity modes.

A new open-source Codex Skill orchestrates voice, avatar, lip-sync, subtitles and QA into a single one-shot pipeline for vertical presenter videos.
No stories match these filters.