In current LLM training, most compute is spent on pre-training, and the reinforcement learning ph..., Sonic AI
“In current LLM training, most compute is spent on pre-training, and the reinforcement learning phase is often stopped early to prevent overfitting to imperfect reward functions.”