The widespread publication of models using Direct Preference Optimization (DPO) creates the appea..., Sonic AI
“The widespread publication of models using Direct Preference Optimization (DPO) creates the appearance that it is the settled best method, but the underlying research questions comparing it to other RL methods remain unanswered.”