“Current reinforcement learning practices in the LLM domain are analogous to the "sim-to-real" approach used in robotics.”