In the "Critiques" project, models trained for helpfulness using RLHF also became significantly b..., Sonic AI
“In the "Critiques" project, models trained for helpfulness using RLHF also became significantly better at finding bugs in code, with an improvement equivalent to a 30x increase in pre-training compute.”