“Current AI models are not learning from visual data to a significant extent and are far from being trained on all available visual inputs.”