“Reinforcement learning is part of the solution for fixing the problem of truthfulness in language models.”