“Most RLHF research papers do not show intermediate checkpoints, which would be valuable for understanding model behavior during training epochs.”