
OpenBMB's MiniCPM5-2B Beats Larger Rivals With 128K Context on Consumer Hardware
OpenBMB's compact 2B model hits open-source SOTA in its class, beating several 4B rivals on code, math, and agent benchmarks while running locally.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

OpenBMB's compact 2B model hits open-source SOTA in its class, beating several 4B rivals on code, math, and agent benchmarks while running locally.

NVIDIA's fine-tuned Nemotron system scored 535.4 out of 600 at IOI 2026, topping the best human contestant under identical contest conditions.
GitHub added a REST endpoint that returns weekly star counts with timestamps, restoring star-history tracking that broke when stargazer listings were locked down.

Artificial Analysis pushes Intelligence Index v4.2 with private test sets, a 4,592-page PDF reasoning benchmark, and a new agentic knowledge-work suite.

Perplexity open-sources details of Ivy, Tulip, and ROSE, its Rust plus Python embedding stack that beats vLLM on latency and throughput.

IFM released six fully open models from 0.9B to 375B parameters, complete with weights, code, training data, and recipes under Apache 2.0.

LlamaIndex and Kaggle launch a schema-guided document extraction benchmark that grades models on missing fields, source grounding, and repeated records across 370 enterprise files.

Claude autonomously wrote a 13 million line Lean proof of Fermat's Last Theorem in 11 days, verifying 29,500 supporting theorems along the way.

A unified pruning, quantization and distillation pipeline shrinks a Vision Transformer 54.5x while holding 95.13% accuracy on out-of-distribution chilli disease images.

GitHub's new research preview routes each coding task through single, cascade, or critique workflows across multiple models, beating Opus 5 at a fraction of the cost.

Google's newest music generation model expands from Flow Music into the Gemini API, AI Studio, and the Gemini app with 44.1 kHz stereo output.

Cohere Labs scraped 696K agent tools from 123K MCP servers and found only 2.6% can actually finish a real job task alone.

InclusionAI open-sources a 124B-parameter mixture-of-experts vision-language model with 5.5B active parameters, 256K context, and MIT license.

A new hybrid Structure-from-Motion framework from NAVER Labs unifies real-time SLAM and offline reconstruction, beating even calibrated systems while running uncalibrated.

A new CLI treats AGENTS.md as a set of weights, mining your agent transcripts to propose evidence-backed edits under a strict token budget.

Anthropic's ant CLI now lets you declare agents, skills, environments, memory stores, and deployments as files and reconcile them like Terraform.

SpaceXAI opens its always-on agent platform to companies with access, network, and audit controls, plus two weeks of free usage for existing customers.

Prime Intellect rebuilt weight sync on RDMA and vLLM tracing, dropping GLM-5.2's 1.6TB policy transfer from 86 seconds to under 4.

Hermes Desktop now auto-detects your hardware, picks a fitting local model, downloads it, and configures the runtime without any manual setup.

OpenAI's new flagship saturates ARC-AGI-3 and FrontierMath Tier 4, drives a browser at 1.9x speed, and crosses the Critical cybersecurity threshold.

Anthropic is proposing TypeScript Function Hooks for Claude Code, letting plugins deeply customize behavior with Express-style middleware and admin-controlled side effects.

Tencent Hunyuan's environment evolution builds off-policy lineages of increasingly hard terminal tasks, boosting Qwen3.6 agents by up to 18 points on Terminal-Bench.

Warp launches Factory Benchmarks, letting teams replay real agent runs to pick the best model per task and cut cost-per-PR from $80 to $30.

Google Research and HHMI Janelia released a wiring diagram of 166,000 neurons and 125 million synapses spanning the male fruit fly brain and nerve cord.
No stories match these filters.