“When a successful machine learning model moves from training to production, the compute workload flips from 100% training to 90-95% inference.”