The DeepSeek team's attempts to use a process reward model (PRM) failed due to the difficulty of ..., Sonic AI
“The DeepSeek team's attempts to use a process reward model (PRM) failed due to the difficulty of defining per-step rubrics and its vulnerability to reward hacking.”