Keep pulling the thread on Oriol Vinyals.
Gemini Robotics-ER 1.5 establishes a new state-of-the-art for embodied reasoning in areas such as visual and spatial understanding, task planning, and progress estimation.
Gemini Robotics 1.5 is a multi-embodiment Vision-Language-Action (VLA) model.
Gemini Robotics 1.5 features a Motion Transfer (MT) mechanism that enables it to learn from heterogeneous, multi-embodiment robot data.
The Motion Transfer (MT) mechanism in Gemini Robotics 1.5 makes the Vision-Language-Action (VLA) model more general.
Gemini Robotics 1.5 interleaves actions with a multi-level internal reasoning process in natural language, enabling the robot to "think before acting".
The internal reasoning process of Gemini Robotics 1.5 improves its ability to decompose and execute complex, multi-step tasks.
The natural language internal reasoning process of Gemini Robotics 1.5 makes the robot's behavior more interpretable to the user.