“A world model learns directly from real-world sensory data such as video, audio, and touch, without relying on language.”