Skip to content
Sonic AI
Larger policy models see less benefit from optimization against a reward model in RLHF and also o..., Sonic AI