“Reinforcement learning allows a mid-tier model to be made as good as a larger-tier model from 3 to 6 months prior.”