On reward modeling tasks, the generalization lines for weak-to-strong learning are mostly flat, i..., Sonic AI
“On reward modeling tasks, the generalization lines for weak-to-strong learning are mostly flat, indicating the strong student model does not significantly outperform the weak teacher model.”