To effectively improve a model over time with PPO, new "on-policy" preference data must be collec..., Sonic AI
“To effectively improve a model over time with PPO, new "on-policy" preference data must be collected from the current version of the instruction-tuned model.”