
RVN Strips Qwen3.8-27B's Refusals Down to Nearly Zero With 6x Better Accuracy
A double-refined abliteration of Qwen3.8-27B cuts refusals to near zero while shrinking behavioral damage roughly sixfold versus its upstream.
hermes-ai.net
Read docs →A focused feed of models, agents, research and open-source releases for people building with AI.

A double-refined abliteration of Qwen3.8-27B cuts refusals to near zero while shrinking behavioral damage roughly sixfold versus its upstream.

A community-built IQ2_XXS quantization squeezes Qwen3.8-Flash-Next's 177B parameters into a 75GB GGUF that runs on ~42GB of memory.

DeepSeek gives its Flash model eyes with an experimental multimodal release that matches Claude Opus 4.8 on agent tasks at a fraction of the cost.

Z.ai shipped a frontier coding model without touching the base weights, matching GPT-5.6 Sol and Claude Fable 5 on agentic tasks through scaled post-training alone.

Apodex 1.1 lands with an Elo of 1348 on GDPval-AA v2, beating DeepSeek V4 Pro and Kimi K2.6 on real-world agentic work.

A community quant of Qwen3.8-Flash-Next shrinks the 177B MoE to 84 GiB with a per-layer mixed-precision recipe that beats standard IQ4_XS on both size and quality.

OrcaRouter stripped the refusal alignment from Z.ai's 320B GLM-5.3-Flash MoE, baking the edit directly into the official block-FP8 shards.

An AI assistant now handles online checkout end to end by connecting to Stripe Link and using single-use virtual cards that need per-purchase approval.

Perplexity Computer now routes long-context, multimodal research tasks to GLM 5.3, which outperformed GLM 5.2 on the in-house WANDR benchmark.

Epoch AI's EBR-bench human baseline shows people quickly outclass frontier models at Earthborne Rangers, exposing a real learning gap.

Tencent's AngelSlim team shrinks the 770B Hy4-preview MoE from 1.5TB down to 213GB using a custom 1.31-bit quantization strategy with minimal accuracy loss.

Singapore lab Sapiens AI pushes its Agnes 2.5 Pro model to 49 on Artificial Analysis Intelligence Index through agentic gains, but at roughly double the token cost.

A community fine-tune of Qwen3.8-27B strips refusals, fixes the resulting weight damage, and ships in four formats tuned for local inference.

Anthropic opens Claude Team seats to 10,000 academic researchers with free Standard access and $15 Premium seats, an 80% cut off list price for a year.

Brett Adcock's stealth AI startup lands gigawatt-scale Vera Rubin capacity, joining Thinking Machines and xAI in NVIDIA's biggest compute alliances.

Google Research unveils a self-supervised model that splits glucose data into slow trends and short spikes, beating prior baselines by 5.8 PR-AUC points.

Google's new speech-to-text model hits 2.6% word error rate, adds screen-aware context in Antigravity, and ships behind two developer APIs.

Z.ai just dropped a 320B mixture-of-experts model with 18B active params, MIT-licensed weights, native multimodal input, and a 1M-token context window.

Alibaba open-weights a 125B multimodal MoE with just 6B active parameters, previewing the attention overhaul coming in Qwen4.

OpenAI adds a $100 middle tier to ChatGPT Business, giving heavy users five times the capacity and killing the five-hour cap.

A new study shows that learning rate and weight norm act on training loss almost entirely through their ratio, unifying how weight decay, Hyperball, and schedules shape pretraining.

Public traces hid 315,320 encrypted blocks. Treat resume logs and publish logs as two paths, then count the field on your own machine.

Z.ai's new 753B open-weights model keeps the GLM-5.2 base but doubles down on post-training, topping open-source coding and cyber benchmarks.

Google Research drops a 330M-parameter foundation model that forecasts multiple related time series jointly, topping three zero-shot benchmarks but locked to non-commercial use.
No stories match these filters.