
Liquid AI's Pipette Exposes What Server Benchmarks Hide About Phone AI
Liquid AI and Artificial Analysis release an open-source suite that measures model quality, speed, latency, and memory across real phones, laptops, and embedded hardware.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Liquid AI and Artificial Analysis release an open-source suite that measures model quality, speed, latency, and memory across real phones, laptops, and embedded hardware.

Jeremy Avigad argues the fixation on neural theorem provers hides a far richer landscape of ways AI is reshaping how mathematics gets done.

A new Zig-built plugin turns x64dbg into an MCP endpoint, letting Claude and other AI agents drive reverse engineering sessions through 71 tools.

Qwen ships an FP8 preview of the architecture behind Qwen4, pairing sparse attention, n-gram embeddings and 125B params with 6B active.

A compact 4B open-source model with hybrid sliding-window attention, native 1M-token context, and agent-focused benchmarks that top comparable small models.

Alibaba's Accio team open sourced a 107-task benchmark that forces agents to complete real e-commerce workflows inside stateful replicas of Shopify, Gmail, Stripe and more.

Alibaba PAI drops a single ControlNet-Union checkpoint that adds Canny, Depth, HED, MLSD, Pose and inpainting to MiniMax-H3 video generation.

A weekend reverse-engineering project turns a compiled macOS app into 440K lines of readable TypeScript, then bolts on a four-provider inference router.

GitHub added a three-day default cooldown on Dependabot version update pull requests, aiming to filter out short-lived poisoned package releases before they land in your repo.

GooeyPi wraps three terminal coding agents in a single Electron desktop app, giving Pi, OMP, and Prime Agent a shared cross platform GUI.

A community quantization pairs a 27B Qwen model with multi-token prediction and AMD's IU4 matrix path, hitting ~49 tokens per second on a single Strix Halo APU.

A community-made checkpoint grafts Z-Image's texture attention onto MiniMax H3's video engine, giving richer surfaces without retraining or extra VRAM.

Anthropic ships reliability fixes to Claude Code Remote Control, adds auto-reconnect, tighter phone-CLI sync, and the ability to launch sessions from your phone.

OpenAI now lets teams attribute spend down to individual API keys and enforce hard monthly caps that cut off traffic when hit.

DeepSeek V4 Pro hits 90.5% on ARC-AGI-1 and 61.3% on ARC-AGI-2, but extra reasoning barely moves the needle on abstract puzzles.

Google Research unveils ME-POIs, a framework that fuses anonymized foot-traffic patterns with text embeddings to improve place understanding by up to 81.9%.

Anthropic opens its most capable cybersecurity model to Enterprise customers through scans, partner integrations, and $35M in open-source credits.

Google's new multi-agent system turns messy wearable sensor streams into statistically vetted biomarker candidates through adversarial validation and human review.

Pika Speech is a 3B flow-matching TTS model that generates a minute of 48 kHz audio in about a second at one cent per minute.

A local router merges Kimi, DeepSeek, and other external models into the Codex picker without kicking out your native GPT models or ChatGPT login.

A new leaderboard scores frontier models on synthesizing 70 to 150 page medical case files, with Claude Fable 5 leading at 64.4 percent.

The open-source document parser Docling has crossed 65,000 GitHub stars, cementing its place as a go-to tool for prepping files for gen AI pipelines.

Rakazo bundles persistent AI teammates with their own live Linux desktops into a self-hosted stack you run on your own hardware and model keys.

SkyRL's IsoExec makes vLLM rollout and Megatron training agree bitwise on token log-probabilities, cutting numerical drift below 1e-6 with 25% overhead.
No stories match these filters.