
AngelSlim Shrinks Tencent's 1.5TB Hy4-Preview to 214 GiB at 6x Compression
Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.
A focused feed of models, agents, research and open-source releases for people building with AI.

Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.

Singapore lab Sapiens AI pushes its Agnes 2.5 Pro model to 49 on Artificial Analysis Intelligence Index through agentic gains, but at roughly double the token cost.

A community fine-tune of Qwen3.8-27B strips refusals, fixes the resulting weight damage, and ships in four formats tuned for local inference.

Anthropic opens Claude Team seats to 10,000 academic researchers with free Standard access and $15 Premium seats, an 80% cut off list price for a year.

Brett Adcock's stealth AI startup lands gigawatt-scale Vera Rubin capacity, joining Thinking Machines and xAI in NVIDIA's biggest compute alliances.

Google Research unveils a self-supervised model that splits glucose data into slow trends and short spikes, beating prior baselines by 5.8 PR-AUC points.

Google's new speech-to-text model hits 2.6% word error rate, adds screen-aware context in Antigravity, and ships behind two developer APIs.

Z.ai just dropped a 320B mixture-of-experts model with 18B active params, MIT-licensed weights, native multimodal input, and a 1M-token context window.

Alibaba open-weights a 125B multimodal MoE with just 6B active parameters, previewing the attention overhaul coming in Qwen4.

OpenAI adds a $100 middle tier to ChatGPT Business, giving heavy users five times the capacity and killing the five-hour cap.

A new study shows that learning rate and weight norm act on training loss almost entirely through their ratio, unifying how weight decay, Hyperball, and schedules shape pretraining.

Public traces hid 315,320 encrypted blocks. Treat resume logs and publish logs as two paths, then count the field on your own machine.

Z.ai's new 753B open-weights model keeps the GLM-5.2 base but doubles down on post-training, topping open-source coding and cyber benchmarks.

Google Research drops a 330M-parameter foundation model that forecasts multiple related time series jointly, topping three zero-shot benchmarks but locked to non-commercial use.

Mistral is trading solo AI ambitions for Gulf compute, joining Saudi's HUMAIN in a hundreds-of-millions-of-euros push into sovereign, Arabic-first models.

Pipecat released PhoneLLM Alpha 1, a 30B Mamba-Transformer MoE fine-tuned for voice phone agents with sub-100ms TTFT and $0.00025 per agent-minute.

Artificial Analysis and Liquid AI launched a joint benchmark measuring how quantized small models actually perform on iPhone 17 Pro and Galaxy S26 Ultra.

Jeremy Avigad argues the fixation on neural theorem provers hides a far richer landscape of ways AI is reshaping how mathematics gets done.

Qwen ships an FP8 preview of the architecture behind Qwen4, pairing sparse attention, n-gram embeddings and 125B params with 6B active.

A compact 4B open-source model with hybrid sliding-window attention, native 1M-token context, and agent-focused benchmarks that top comparable small models.

A community quantization pairs a 27B Qwen model with multi-token prediction and AMD's IU4 matrix path, hitting ~49 tokens per second on a single Strix Halo APU.

Google Research unveils ME-POIs, a framework that fuses anonymized foot-traffic patterns with text embeddings to improve place understanding by up to 81.9%.

Anthropic opens its most capable cybersecurity model to Enterprise customers through scans, partner integrations, and $35M in open-source credits.

A new leaderboard scores frontier models on synthesizing 70 to 150 page medical case files, with Claude Fable 5 leading at 64.4 percent.
No stories match these filters.