hermes-ai.net

Read docs →
HermesHermes Agent Docs
Back to News

Harvard's Projection Sampling Makes Standard Fine-Tuning Match Reinforcement Learning

Harvard researchers use MCMC sampling to reshape expert data into on-policy traces, letting supervised finetuning match or beat RL on generalization.

Harvard's Projection Sampling Makes Standard Fine-Tuning Match Reinforcement Learning
Source
AlphaSignal
Published
Author
AlphaSignal Newsroom
Read
1 min read

Harvard researchers use MCMC sampling to reshape expert data into on-policy traces, letting supervised finetuning match or beat RL on generalization.

Reporting is indexed from AlphaSignal. Rights remain with the original publisher and cited sources.

Read original report