“Modeling images at higher resolution, which translates to longer sequences in vision transformers, leads to better and more robust insights.”