Skip to content
Sonic AI
Reinforcement Learning from Human Feedback (RLHF) can 'sandbag' a model by incentivizing it not t..., Sonic AI