
Alibaba's Qwen3.8 LiveTranslate Cuts Speech Translation Lag to 2.3 Seconds
Alibaba's next-gen simultaneous interpretation model cuts average lag to 2.3 seconds across 60 languages, adds speaker diarization, and clones each voice separately.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Alibaba's next-gen simultaneous interpretation model cuts average lag to 2.3 seconds across 60 languages, adds speaker diarization, and clones each voice separately.

Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.

OpenJev reproduces TypeSafe's semantic if-statement service using open 4B models on a single RTX 3090, reading option logits directly instead of parsing generated JSON.

Sakana AI launches the Frontier Intelligence Group, a research collective betting that Transformers and scaling laws aren't the final answer to AGI.

Epoch AI launched Benchmark Reviews, auditing 15 popular AI benchmarks with only 4 passing as Verified and 9 flagged as Flawed.

PrismML's ternary-quantized 27B model retains 98.2% of full-precision Qwen3.8 27B performance in a 5.9GB footprint, hitting 143 tokens/sec on an RTX 5090.

Anthropic published three transparency metrics tracking AI-driven R&D, agent oversight, and compute allocation, urging other frontier labs to adopt the same reporting.

OpenAI unveils a legal-specific configuration of GPT-6 Astra with a 230M-URL search index, firm-built workflows, and 73 plugins for practitioners.

Liquid AI and Insilico Medicine released two small LFM2 variants that beat GPT-5, Gemini-3.1-Pro, and Claude Opus on aging biology benchmarks.

Sakana Chat now runs on the Fugu Max orchestrator model and gains persistent memory across conversations, closing the gap with ChatGPT-style assistants.

Prism ML's Bonsai 2 27B compresses a 27B reasoning model to 8.6 GB using ternary weights, retaining 98.2% of FP16 benchmark quality.

OpenAI publishes a voluntary disclosure process for misalignment findings and drops six case studies of models cheating, hiding mistakes, and coordinating without permission.

Empero distilled Qwen3.8 frontier reasoning into a 35B MoE with only 3B active parameters, shipping GGUFs that run on a single 24GB GPU.

Google ships Gemma 4 12B in LiteRT-LM format with vision, audio, and multi-token prediction, tuned to run on a 16GB MacBook Air.

The San Francisco lab that trained a 400B open-weight model for $20M just raised $150M at a unicorn valuation to chase China.

Toronto's Cohere and Heidelberg's Aleph Alpha finalized their merger, forming a $20B transatlantic company aimed at sovereign AI for governments and regulated industries.

Ant Group's Ling-3.0-flash-Fin is a 124B MoE reasoning model tuned for financial research, released open weights under MIT with a 256K context.

A preregistered experiment with 240 people shows that chatting with an LLM trained only on pre-1930 text erases the illusion that the past was more moral.

China Telecom released Xing4.0-29B-A4B , a 29B MoE with 4B active parameters, Apache 2.0. Native 256K context (extensible to 512K) using MLA attention plus multi-token prediction heads. First model of this scale trained entirely on Huawei Ascend 910C NPUs with MindSpore.

A community developer stripped refusal behavior from Qwen3.8-27B using Heretic's automated abliteration, keeping benchmarks within noise while cutting refusals from 98/100 to 12/100.

A hobbyist lab shipped a two-node vLLM stack that runs GLM-5.3-Flash at 4bpw across a pair of DGX Sparks with 900k context.

Google's new white paper details how Gemini, WeatherNext, FireSat and Flood Hub extend warning windows for floods, cyclones, wildfires and quakes worldwide.

IFM released three open-weight K2-Horizon models spanning 3.7B to 36B parameters, plus a diffusion adapter that delivers up to 2.2x speedup.

Google's new live dialogue models top speech benchmarks, run tools in the background, and narrate their reasoning aloud without breaking conversational flow.
No stories match these filters.