
MATLOWAI Rebuilds MiniMax H3 to Run Video and Audio on 16GB Cards
A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

A 21 GB single-file build of MiniMax-H3 renders 10-second clips with synchronized stereo audio in about 76 seconds at 4 steps.

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

OrcaRouter stripped the refusal alignment from Z.ai's 320B GLM-5.3-Flash MoE, baking the edit directly into the official block-FP8 shards.

OpenAI will cut off Cursor's direct model access on November 12, citing distrust of SpaceX after the Musk-led firm's $60B takeover.

A UChicago team shows that fine-tuning models into 'evil' behavior is not mysterious, but predictable from activation geometry across 12 model-dataset setups.

Perplexity's new Search API takes the top three spots on the Artificial Analysis Search Index, extending the quality-cost Pareto frontier for agentic search.

An AI assistant now handles online checkout end to end by connecting to Stripe Link and using single-use virtual cards that need per-purchase approval.

Perplexity Computer now routes long-context, multimodal research tasks to GLM 5.3, which outperformed GLM 5.2 on the in-house WANDR benchmark.

Epoch AI's EBR-bench human baseline shows people quickly outclass frontier models at Earthborne Rangers, exposing a real learning gap.

Exa's new search primitive treats tokens as the retrieval unit rather than documents, cutting agent context bloat by an average of 95%.

Anthropic gave Claude 48 hours and one GPU to fix ten alignment failures in small models, then tested if it could align a frontier successor.

GitHub Copilot lands in Slack and Teams as a shared agent, adds a Customize hub, new models, CLI session sidebar, and on-device dictation.

Krea previews a next-generation foundation model with strong editing plus a dedicated agent platform for creative work with MCP and API hooks.

Sesame releases an open benchmark that scores voice agents on when they speak, yield, or stay silent, exposing where every current system fails.

Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.

Sarvam opens its multilingual voice and video workspace to the public, bundling dubbing, cloning, and TTS across 11 Indian languages into one browser tool.

Midjourney is testing a V8.2 edit model that handles instruction-based editing, multi-image composition, inpainting, and outpainting in one system.

A community port strips NVIDIA's text-to-motion diffusion model down to C++ and GGML, running SMPL-X skeleton generation on CPU or Vulkan.

A free, framework-free Colab curriculum walks backend engineers through the applied LLM stack, from raw API calls to serving, evals, agents, and red-team benchmarks.

Kimi Code 0.39.0 ships an experimental Remote Control mode that lets you drive a local coding session from any browser or phone.

A 26B multimodal Gemma variant with refusal directions surgically removed lands on Hugging Face in GGUF format, ready for llama.cpp.

Sapient Intelligence open-sourced PRAXIST, an orchestration system that turns any runnable project with a measurable metric into a self-directed research loop.

Singapore lab Sapiens AI pushes its Agnes 2.5 Pro model to 49 on Artificial Analysis Intelligence Index through agentic gains, but at roughly double the token cost.

Chroma unveils Fission, a concurrency protocol for agent swarms that skips rollbacks to preserve expensive reasoning tokens when writes collide.
No stories match these filters.