
Tencent Releases Octop, a Self-Hosted Multi-User AI Platform in One Process
Tencent Cloud open-sourced Octop, a single-process, self-hosted AI assistant that gives every user their own team of MBTI-flavored agents across web, CLI, and chat apps.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

Tencent Cloud open-sourced Octop, a single-process, self-hosted AI assistant that gives every user their own team of MBTI-flavored agents across web, CLI, and chat apps.

Researchers isolated a linear direction in 25 open-weight LLMs that behaves like pain, distinct from fear or sadness, and drives self-relief behavior.

xAI's new image model jumps 14 spots to fourth place on the Artificial Analysis leaderboard, undercutting OpenAI's flagship on price by more than 3x.

Anthropic taps Accenture's Faculty unit to embed evaluators inside its labs, with each side pledging $1B over five years to build safety oversight capacity.

Vals AI's new CUA-Bench pits frontier models against six commercial video games with only pixels in and keystrokes out, and every model scores below 20%.

SpaceXAI's new speech-to-text model doubles accuracy over v1.0, tops the streaming leaderboard, and keeps pricing at ten cents per hour.

Factory rolls out on-premises and air-gapped deployment for its Droid coding agents, targeting enterprises that cannot ship source code to a vendor cloud.

Bolt adds DeepSeek V4.1 Flash to its Forge open-model lineup, promising 10x the usage of V4 Pro and stronger visual design output.

Epoch AI marks the first Major Advance on its unsolved-math benchmark, with GPT-6 Astra driving the proof in an interactive session with three mathematicians.

Meta engineers built a domain expert AI that separates knowledge from reasoning and learns from expert corrections without any model retraining.

OpenJev reproduces TypeSafe's semantic if-statement service using open 4B models on a single RTX 3090, reading option logits directly instead of parsing generated JSON.

A reverse-engineered toolkit lets AI agents drive CapCut's Chinese sibling Jianying entirely from the command line, producing native editable drafts and MP4 exports.

Pydantic AI now supports Jev, a non-generative classifier that answers typed questions in about 180 ms, turning your output_type into the question itself.

A new open-source Claude Code plugin uses TypeSafe's Jev decision model to prune agent context per tool call instead of summarizing it away.

Shanghai AI Lab open-sourced Atria Dawn Preview, a 744B-parameter Mixture-of-Experts agentic model with MIT-licensed weights and FP8 checkpoints.

Alibaba's new omni-modal model reasons over audio and video, orchestrates tools across long workflows, and cuts video input costs by roughly 89%.

Sakana AI launches the Frontier Intelligence Group, a research collective betting that Transformers and scaling laws aren't the final answer to AGI.

A new preprint reframes recursive self-improvement as dreaming inside a replay simulator built from an agent's own discovery history.

Factory's coding agent joins Slack's new multiplayer channels in beta, letting teams delegate software tasks and get results delivered inside the conversation.

OpenAI wraps GPT-6 Astra in a legal search index of 230 million URLs, hits 54% correctness on Vals AI's Legal Research Bench.

LM Studio's Bionic agent gains the ability to search its own past transcripts, recovering context lost during compaction and pulling in other sessions with @ mentions.

LM Studio's Bionic agent gains tools to search its own past transcripts, recovering details lost to context compaction and letting users @-mention old sessions.

Epoch AI launched Benchmark Reviews, auditing 15 popular AI benchmarks with only 4 passing as Verified and 9 flagged as Flawed.

Claude accelerated over 30 open-source biology models by roughly 4x, letting a single GPU node predict structures once out of reach.
No stories match these filters.