Reinforcement learning on datasets with automatically checkable solutions, such as STEM problems ..., Sonic AI
“Reinforcement learning on datasets with automatically checkable solutions, such as STEM problems or coding tasks, can significantly improve Chain-of-Thought reasoning capabilities.”