
Midjourney's V8.2 Edit Model Merges Inpainting, References and Instructions Into One
Midjourney is testing a V8.2 edit model that handles instruction-based editing, multi-image composition, inpainting, and outpainting in one system.
A focused feed of models, agents, research and open-source releases for people building with AI.

Midjourney is testing a V8.2 edit model that handles instruction-based editing, multi-image composition, inpainting, and outpainting in one system.

A 26B multimodal Gemma variant with refusal directions surgically removed lands on Hugging Face in GGUF format, ready for llama.cpp.

Google Research unveils an autonomous AI system that turns natural-language questions about food security, disease, and climate risk into trained geospatial models in minutes.

A new keypoint detector skips deblurring entirely, learning directly from blurred images through self-supervision and beating supervised baselines on matching and localization.

A group of vision researchers argues that pure vision, not language-tethered multimodal models, could be its own route to general intelligence.

A new agent skill turns a single reference image into diffable TypeScript that procedurally reconstructs the object as an animation-ready Three.js scene.

Alibaba's third-generation image model debuts at #6 in editing and #9 in text-to-image, with big Elo jumps and a productivity-first pitch.

GPT-Image-2 can now generate PNGs with real alpha channels directly through the API, skipping the separate background-removal step for cutouts.

Meta shared new benchmarks, robotics demos, and a real-world agent evaluation for Muse Spark 1.2, its coding-focused multimodal model ahead of an open-weights release.
No stories match these filters.