
DeepSeek's V4-Flash-Vision-Exp Quietly Challenges Anthropic's Opus on Multimodal Agent Tasks
DeepSeek's new experimental vision model bolts image understanding onto V4-Flash at the same price, closing in on Claude Opus 4.8 on multimodal agent tasks.
A focused feed of models, agents, research and open-source releases for people building with AI.

DeepSeek's new experimental vision model bolts image understanding onto V4-Flash at the same price, closing in on Claude Opus 4.8 on multimodal agent tasks.

Sakana AI upgraded its free JP-EN-ZH translator to the new Namazu model, beating Google Translate, DeepL, and Claude Opus 4.8 in head-to-head evaluations.

Google's mid-tier reasoning model nearly clears the 85% target on ARC-AGI-2 at a quarter per task, redrawing the cost-performance frontier.

Liquid AI ships DSpark draft models for its LFM2.5 family, delivering up to 3.18x GPU throughput and cutting agent latency by nearly half.

Kakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.
No stories match these filters.