“OpenAI's paper "Let's Verify Step-by-Step" is a significant work on process reward models for chain-of-thought reasoning.”