
Ant Group's Ling-3.0-flash-VL Adds Vision to a 124B Open Reasoning Model
Ant Group's InclusionAI lab dropped an MIT-licensed 124B mixture-of-experts vision model that activates just 5.5B parameters per token and ingests text, images, and video.
A focused feed of models, agents, research and open-source releases for people building with AI.

Ant Group's InclusionAI lab dropped an MIT-licensed 124B mixture-of-experts vision model that activates just 5.5B parameters per token and ingests text, images, and video.

OpenAI says GPT-6 Astra needs leaner skills, trimmed AGENTS.md files, and clearer completion criteria to avoid wasted context and premature stops.
PyTorch 2.14 lands with a new CUTLASS-based GEMM backend, in-place fault tolerance for distributed training, and native linear algebra on Apple Silicon.

Google's agentic IDE ships a bundle of upgrades spanning genomics skills, deep-reasoning multi-agent workflows, generative UI artifacts, and a rebuilt terminal.

Cognition brings its two-agent Fusion harness to Devin CLI and Desktop, pairing a frontier planner with a cheap executor for up to 39% savings.

Google's WikiSkill paper gained 15 points of accuracy from a notebook its agent couldn't read, and its skills still need a transfer test

NVIDIA Labs released SoL-Pi, a coding agent extension that cuts token cost by roughly one-third using four mechanisms discovered by auto-research loops.

Comfy-Org has repackaged the MAP team's YuE2 music model into single-file safetensors, and ComfyUI now ships native nodes to run it.

Sakana AI's Fugu Max and Fugu Ultra v2 orchestrate pools of open models to beat frontier LLMs on cost and capability simultaneously.

Google Research introduces ToolGrad, an answer-first framework that generates tool-use training data with a 99.8% pass rate at lower cost.

Cursor's new Projects feature replaces one-off chats with a persistent coordinator agent that delegates to subagents, syncs context, and runs recurring work.

NVIDIA and USC present HorizonRelight, a diffusion transformer method that keeps lighting stable across long videos by propagating context between chunks.

OpenAI opens up the same agent harness that powers Codex, letting developers spin up long-running cloud agents with a single API call.

Anthropic ships a live session viewer and a server-side auto permission mode for Claude Managed Agents, letting the model decide when to pause for human approval.

OpenAI launches a vertical ChatGPT for banks that bundles Daloopa, PitchBook, and LSEG data with GPT-6 Astra reasoning and firm-specific templates.

humans& releases Persimmon, a 550B user model built to simulate real people in group chats, fooling AI judges 20 percent of the time.

Runway details how it turns diffusion video models into causal, autoregressive streamers using teacher forcing and on-policy distillation.

Together AI ported its ThunderKittens kernel framework to NVIDIA's Vera Rubin NVL72, hitting 22.4 PFLOPS on NVFP4 GEMMs and rivaling cuBLAS.

A math benchmark built to resist AI just fell. GPT-6 Astra cracked the final Tier 4 problem, closing out a 98 percent run in 14 months.

OpenAI's full-duplex voice model lands in the API at $0.05 per minute, letting agents listen and speak simultaneously while delegating reasoning to backend models.

Anthropic's latest threat report details how state actors, criminals, and hacktivists weaponized Claude across cyber, surveillance, biology, and distillation campaigns.

A new recurrent reasoning method mixes denoising with looped hidden states, hitting 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 with just 7M parameters.

Edge0 is an open source framework that streams MoE experts from SSD, letting a 35B parameter model run on a 24GB Mac mini with under 3GB active memory.

Google folded the full Gemini API documentation directly into AI Studio, so developers can read reference material without leaving the build surface.
No stories match these filters.