Standard RL training for correctness on the GSM dataset caused human judge accuracy to decrease o..., Sonic AI
“Standard RL training for correctness on the GSM dataset caused human judge accuracy to decrease over time, in contrast to the improvement seen with "checkability training".”