“The three-stage SFT-DPO-RL recipe is no longer sufficient for creating frontier reasoning and agentic models.”