“Current vision-language models will not achieve the level of world understanding of a cat or a dog.”