
OpenBMB's MiniCPM5-2B Tops Sub-4B Charts While Burning 3x Fewer Tokens
OpenBMB's 2.6B dense reasoning model tops the sub-4B open weights leaderboard, punching well above its weight on agentic benchmarks while staying token-efficient.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

OpenBMB's 2.6B dense reasoning model tops the sub-4B open weights leaderboard, punching well above its weight on agentic benchmarks while staying token-efficient.

A year-long Stanford study of 1,182 Character.AI users finds heavier chatbot engagement predicts lower well-being, largely by crowding out face-to-face contact.

TokenRhythm fine-tunes Qwen3.5-4B with a routing harness that turns agent traces into training data, gaining 5.93 points across ten benchmarks.

Japan's national LLM project releases an Apache-2.0 vision-language model with reasoning traces, trained on a scrubbed 29M-sample dataset.

OpenBMB's compact 2B model hits open-source SOTA in its class, beating several 4B rivals on code, math, and agent benchmarks while running locally.

Artificial Analysis pushes Intelligence Index v4.2 with private test sets, a 4,592-page PDF reasoning benchmark, and a new agentic knowledge-work suite.

IFM released six fully open models from 0.9B to 375B parameters, complete with weights, code, training data, and recipes under Apache 2.0.

InclusionAI open-sources a 124B-parameter mixture-of-experts vision-language model with 5.5B active parameters, 256K context, and MIT license.

OpenAI's new flagship saturates ARC-AGI-3 and FrontierMath Tier 4, drives a browser at 1.9x speed, and crosses the Critical cybersecurity threshold.

Keep the project on disk, treat lunch as a bill, and score the first finished tasks rather than dollars per million tokens

Google DeepMind's new weather AI trains on raw satellite feeds and station data, delivering hourly 5km forecasts with up to 50% better rain accuracy.

Xiaohongshu's Self-GC uses a planner LLM to fold, mask, or prune agent context, retaining critical details 84.85% of the time versus 54.55% for rule-based baselines.

An 8-year interpretability project shows that LLM representations can be closely approximated by symbolic role-filler structures, enabling precise behavioral edits.

Google's third Flash release in as many months pushes a workhorse model into frontier-tier territory on agentic coding, legal, and finance benchmarks.

Multiverse Computing's Quasar 438B tops European AI rankings with a 43 Intelligence Index score and 15.3-second reasoning responses.

Alibaba's flagship 2.4T-parameter model gets a coding and agent-work refresh with the same pricing and 1M context window.

Independent evaluations put Anthropic's new flagship at the top of the intelligence charts, but the win comes with a token-count tax that eats the cache savings.

A new benchmark study finds that sliding-window attention with sinks matches or beats post-trained linear attention models, without any retraining.

Cursor added Anthropic's new Fable 5.1 model, which posted a 73.4% score on CursorBench 3.2 and excels at self-verifying long coding runs.

Shieldstral accepts policies at runtime. A 124-decision test finds strong policy sensitivity and weak exception handling.

New scaling laws show that looping the middle half of a Mixture-of-Experts model twice saves up to 18% of training compute at matched budgets.

A community fine-tune of Qwen3.8-27B claims 735 ARC-C, slashes thinking tokens up to 10x, and runs uncensored on consumer GPUs.

Google's new 330M-parameter time series foundation model handles multivariate forecasting zero-shot in a single forward pass, topping Gift-Eval, FEV-Bench and Time.

Seven practical shifts for securing agents that can find any crack
No stories match these filters.