“The scaling constraints for the o1 model's training approach are substantially different from those of standard LLM pretraining.”