←Back to NewsAI News/Llmsdeep-diveLlmsInfraWhat DeepSeek-V4.1-Flash teaches us about efficient AIA 552B model with 890-byte KV cache, 8B active on input, and an Artificial Analysis 40 versus Gemini 3.8 Flash High at 41SourceBen DicksonPublishedSep 14, 2026, 4:00 PMAuthorAlphaSignal NewsroomRead1 min readA 552B model with 890-byte KV cache, 8B active on input, and an Artificial Analysis 40 versus Gemini 3.8 Flash High at 41Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsAlphaSignal · repoMia AI Lab Runs GLM-5.3-Flash Across Two DGX Sparks at 146 tok/sGoogle DeepMind · newsGoogle DeepMind's WeatherNext 3 Ditches Physics Simulations for 60% Sharper ForecastsAlphaSignal · paperMicrosoft Finds a 2023 Trick Beats Two Years of Linear Attention Research