“With proper training, jointly training vision and text modalities can be mutually enhancing, improving performance in both domains.”