
World Labs' Atlas Unifies 3D Reconstruction and Video Generation in One Model
World Labs unveiled Atlas, an omni world model that unifies text, images, video, and 3D into a single 3D-grounded spatial context.
A focused feed of models, agents, research and open-source releases for people building with AI.

World Labs unveiled Atlas, an omni world model that unifies text, images, video, and 3D into a single 3D-grounded spatial context.

Krea's new Realtime Director lets you steer a live AI video stream with prompts as it plays, powered by fal's H3 Max Director model.

Unitree's UnifoLM-X2-1.0 lets humanoid robots spar autonomously by predicting the next few seconds of physics in real time.

VDN-H3 replaces most of MiniMax H3's quadratic attention with a linear branch, rendering a 14.4-second 768p clip in 11.23 seconds on 8 B200s.

A community LoRA for MiniMax H3 sharpens video output through aligned guide latents, offering a new pass on ComfyUI's ref2va pipeline.

Runway's new general world model streams continuous 720p video and 48kHz audio in real time, taking text prompts and camera input as you play.

A ComfyUI custom node makes MiniMax H3 clips chain seamlessly by carrying motion and the exact audio waveform across cuts.

Amap CV Lab open-sourced ABot-Recon, a streaming 3D reconstruction model that runs on a rolling 12-frame window at 24 FPS.

A new open-source agent skill hands Claude Code and Codex a full motion-design studio for voiceover explainer videos with word-level sync and Remotion rendering.

Utopia is a self-hosted Rust and Postgres knowledge platform that layers RAG on top of a bitemporal knowledge graph tracking when each fact was true.

Runway Ruby now exports scene-referred half-float EXR sequences in ACEScg 1.3 and 2.0, slotting AI video directly into professional color pipelines.

World Labs unveiled Atlas, an omni world model that generates 1440p video with pixel-perfect camera control and reconstructs scenes in 3D from a handful of images.

JD.com open sources a 16B autoregressive diffusion editor that rewrites live video at 30 FPS in 720p from a text prompt.

Pireel is an open-source, browser-based AI video editor for talking-head content that any AI agent can drive over MCP.

Runway's Solaris skips the code layer entirely, generating live app interfaces frame by frame and beating top LLMs on visual reconstruction tests.

A new Gaussian splatting framework reconstructs volumetric videos with thousands of frames of complex motion, extending prior methods by roughly 70x in temporal coverage.

A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.

Sarvam opens its multilingual voice and video workspace to the public, bundling dubbing, cloning, and TTS across 11 Indian languages into one browser tool.

Google's video model gets scene extension to 40 seconds, keyframe control, 360p drafting, 4K upscaling, and video reference inputs.

Krea launches MiniMax H3 Max, a speed-tuned variant of the Hailuo 3 video model that renders 15 second clips in roughly 5 seconds.

A community developer shipped a 54-node ComfyUI extension for MiniMax H3 with same-day tooling for Alibaba PAI's new 8-step acceleration LoRAs.

FastVideo distilled MiniMax's 33B video-and-audio diffusion transformer from 50 denoising steps down to four, with open weights and 90% sparse attention.

A group of vision researchers argues that pure vision, not language-tethered multimodal models, could be its own route to general intelligence.

Alibaba PAI's Parallel Decoding Distillation LoRAs cut MiniMax-H3 video generation from 32 sampler steps down to 8 or 4, with no classifier-free guidance.
No stories match these filters.