“Increasing test-time compute, allowing models to "think" for up to 30 minutes, has yielded great returns in performance.”