“The GROOT-N1 and N1.5 models use a Vision Language Model (VLM) backbone and a diffusion transformer to output a float vector for robot actions.”