“Reinforcement learning should be used minimally because it is extremely inefficient in terms of sample efficiency.”