Skip to content
Sonic
AI
Sonic
AI
Home
Discover
Library
Ask Sonic
Projects
Use with Claude or ChatGPT
Show me around
Request source or feature
Using more reward model data in RLHF leads to higher gold reward scores and reduces "Goodharting"., Sonic AI
Back
“Using more reward model data in RLHF leads to higher gold reward scores and reduces "Goodharting".”
Lilian Weng
LLMs
Loading full analysis…