In experiments across four RL environments, models with higher capability (larger size, higher ac..., Sonic AI
“In experiments across four RL environments, models with higher capability (larger size, higher action resolution, better observation fidelity, more training) tend to achieve higher proxy rewards but decreased true rewards.”