
xAI Ships Grok Voice Transcribe 2.0 With Half the Errors at Same Price
SpaceXAI's new speech-to-text model doubles accuracy over v1.0, tops the streaming leaderboard, and keeps pricing at ten cents per hour.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

SpaceXAI's new speech-to-text model doubles accuracy over v1.0, tops the streaming leaderboard, and keeps pricing at ten cents per hour.

OpenAI unveils a legal-specific configuration of GPT-6 Astra with a 230M-URL search index, firm-built workflows, and 73 plugins for practitioners.

Exa Snapshot pins search and page retrieval to any past date, using 400 billion stored webpages to prevent evaluation leakage and enable backtests.

Claude Code projects now run parallel cloud threads coordinated by a chief-of-staff style controller, with shared memory that persists across sessions.

Runway's new Enhance Frame Rate model retimes video to any target between 24 and 120 fps, including broadcast fractional rates, priced by input seconds.

Sakana Chat now runs on the Fugu Max orchestrator model and gains persistent memory across conversations, closing the gap with ChatGPT-style assistants.

Grok Bot now plugs into 1Password so its cloud browser can log into any site without your agent ever seeing the raw password.

Sakana Marlin adds Interactive Reading for source-grounded chat with reports and editable PowerPoint export with custom templates.

Google's new live dialogue models top speech benchmarks, run tools in the background, and narrate their reasoning aloud without breaking conversational flow.

Anthropic's new plugin drops live Salesforce pipelines, accounts, and 37 prebuilt sales skills into Claude, letting sellers act on CRM data without leaving the chat.

OpenAI's new life sciences reasoning model moves from research preview to trusted access, with a Codex plugin connecting to 50+ scientific databases and tools.

OpenAI opens up the same agent harness that powers Codex, letting developers spin up long-running cloud agents with a single API call.

OpenAI's full-duplex voice model lands in the API at $0.05 per minute, letting agents listen and speak simultaneously while delegating reasoning to backend models.

Google folded the full Gemini API documentation directly into AI Studio, so developers can read reference material without leaving the build surface.

Google launches a native Gemini desktop app for Windows 10 and 11 with an Alt+Space hotkey, Spark agent access, and image and video generation.

Google Pics packages Nano Banana into a Workspace-native image editor with object-level control, in-image text editing, and live collaboration.

Black Forest Labs shipped a prompt-driven video editor that changes one thing per clip while leaving length, framing, and audio untouched.

DeepSeek's 552B mixture-of-experts model splits input and output compute asymmetrically, natively handles vision, and undercuts its own flagship on price.

OpenAI's new image model brings 50% faster generation, precision comment-based edits, in-chat sketching, and two new API tiers for developers.

Inception's new diffusion-based LLM hits 1,107 tokens per second on NVIDIA GPUs while boosting quality 40% over Mercury 2.

Gradium's new Voice Design turns a written description into a brand-new synthetic voice in seconds, no cloning or licensing required.

DeepMind precomputed AlphaGenome predictions for all 9 billion single-letter human DNA mutations, delivering a 1-petabyte searchable atlas with impact scores.
GitHub added a REST endpoint that returns weekly star counts with timestamps, restoring star-history tracking that broke when stargazer listings were locked down.

IFM released six fully open models from 0.9B to 375B parameters, complete with weights, code, training data, and recipes under Apache 2.0.
No stories match these filters.