“There is evidence from the last 5-6 years of clean log-linear scaling with compute in reinforcement learning.”