Keep pulling the thread on Oriol Vinyals.
A new family of AI models named Gemini Robotics has been developed for robotics applications, built upon the foundation of Gemini 2.0.
Gemini Robotics is an advanced Vision-Language-Action (VLA) generalist model capable of directly controlling robots.
The Gemini Robotics model can be fine-tuned to learn new short-horizon tasks from as few as 100 demonstrations.
The Gemini Robotics model can be fine-tuned to adapt to completely novel robot embodiments.
The Gemini Robotics model can execute smooth and reactive movements to perform a wide range of complex manipulation tasks.
The Gemini Robotics model is robust to variations in object types and positions, can handle unseen environments, and can follow diverse, open-vocabulary instructions.
With additional fine-tuning, the Gemini Robotics model can be specialized to solve long-horizon, highly dexterous tasks.
The Gemini Robotics-ER (Embodied Reasoning) model extends Gemini's multimodal reasoning capabilities into the physical world with enhanced spatial and temporal understanding.
The Gemini Robotics-ER model's capabilities include object detection, pointing, trajectory and grasp prediction, multi-view correspondence, and 3D bounding box predictions.