The Llama 3 recipe does not use online Reinforcement Learning; instead, a reward model filters sa..., Sonic AI
“The Llama 3 recipe does not use online Reinforcement Learning; instead, a reward model filters samples over six rounds, with the best models from each round seeding the next.”