Skip to content
Sonic
AI
Sonic
AI
Home
Discover
Library
Ask Sonic
Projects
Use with Claude or ChatGPT
Show me around
Request source or feature
A reward model was found to resist human preference on over 25% of the HHH-RLHF dataset., Sonic AI
Back
“A reward model was found to resist human preference on over 25% of the HHH-RLHF dataset.”
Lilian Weng
LLMs
Loading full analysis…