“Having models produce a detailed inner monologue allows for shorter-horizon feedback during RL training, which makes the system safer.”