←Back to NewsAI News/Training InfrapaperTraining InfraLlmsKakao Trains a 155B MoE on 10T Tokens Without a Single Hyperparameter SweepKakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.SourceAlphaSignalPublishedAug 20, 2026, 1:57 PMAuthorAlphaSignal NewsroomRead1 min readKakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.Read original report ↗Next readsMagic · newsMagic Matches DeepSeek V4 Pro Base Using 50x Less ComputeAlphaSignal · paperTsinghua's SMELT Cuts AI Training Costs 18% by Looping Layers TwiceHark · newsFigure AI's Brett Adcock Bets Gigawatt NVIDIA Deal on Hark Handoff