
audio.cpp Runs Speech and Music AI up to 5x Faster Without Python
A pure C++ inference engine built on ggml runs TTS, ASR, voice cloning, and music generation up to 5x faster than Python baselines with no Python dependency.
A focused feed of models, agents, research and open-source releases for people building with AI.

A pure C++ inference engine built on ggml runs TTS, ASR, voice cloning, and music generation up to 5x faster than Python baselines with no Python dependency.

A Tsinghua team trained a 2B language model from scratch on consumer RTX 5090s for under $6,900, matching Qwen2.5-1.5B and open-sourcing everything.

JD.com open sources a 16B autoregressive diffusion editor that rewrites live video at 30 FPS in 720p from a text prompt.

Pireel is an open-source, browser-based AI video editor for talking-head content that any AI agent can drive over MCP.

Pollen Robotics open-sourced Microduck, a 25 cm biped robot brain written in Rust that walks via reinforcement learning policies exported to ONNX.

Perplexity's Computer now splits agent tasks between cloud frontier models and an on-device model, gated by an open-source 0.6B PII detector.

EvoMap's AutoResearch is an open-source agent workflow that takes a research idea from discovery through experiments to a paper-ready evidence package.

LM Studio's agentic app Bionic now runs on Linux, giving penguin users repo-aware coding, document work, and voice transcription with local open models.

A double-refined abliteration of Qwen3.8-27B cuts refusals to near zero while shrinking behavioral damage roughly sixfold versus its upstream.

A community port strips NVIDIA's text-to-motion diffusion model down to C++ and GGML, running SMPL-X skeleton generation on CPU or Vulkan.

A new keypoint detector skips deblurring entirely, learning directly from blurred images through self-supervision and beating supervised baselines on matching and localization.

A local Python CLI parses session logs from Claude Code, Codex, and Gemini CLI to tally token usage and spend by model, project, and day.

QoderAI's Better Harness plugs into Claude Code, Codex, Cursor and Copilot to audit how your coding agents actually work, not just what they ship.

FastVideo distilled MiniMax's 33B video-and-audio diffusion transformer from 50 denoising steps down to four, with open weights and 90% sparse attention.

SenteLabs open-sourced an eight-agent Claude system that simulates a full C-suite, complete with episodic memory, RAG, and a persistent executive voice.

Zed 1.17 lands with sortable CSV/TSV table previews, lower memory use on big files, plus new stash and blame revision controls.

Z.ai just dropped a 320B mixture-of-experts model with 18B active params, MIT-licensed weights, native multimodal input, and a 1M-token context window.

Alibaba open-weights a 125B multimodal MoE with just 6B active parameters, previewing the attention overhaul coming in Qwen4.

Breeze TTS 2 tops the open weights speech leaderboard with a 90 Elo lead, combining voice design, direction, and sub 140ms streaming.

Graft builds a persistent, plain-English graph of your codebase so coding agents skip re-exploration and hit 66% on SWE-bench Verified.

BreezeBlue open-sourced a 3B bilingual TTS model that tops the Artificial Analysis leaderboard with sub-40ms latency and natural-language voice design.

A new open-source Rust tool renders Claude Code sessions as a live, scrubbable flow graph inside your terminal or browser.

Apodex open-sourced FrontierAgent, an Apache 2.0 terminal agent framework with ReAct and Agent Team modes, alongside a 35B open-weight model.

Kyutai released the full training stack for its 100M-parameter Pocket TTS, letting anyone train a voice model from scratch for under $200.
No stories match these filters.