hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Kakao Trains a 155B MoE on 10T Tokens Without a Single Hyperparameter Sweep

Kakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.

Kakao Trains a 155B MoE on 10T Tokens Without a Single Hyperparameter Sweep
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

Kakao researchers show how a two-step transfer trick predicts the optimal learning rate for a 10-trillion-token MoE run using tiny proxy models.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report