Skip to content
Sonic AI
On reward modeling tasks, the generalization lines for weak-to-strong learning are mostly flat, i..., Sonic AI