Keep pulling the thread on Sergey Levine.
Physical Intelligence is developing robotic foundation models with the goal of enabling any physical robot to perform any task in any environment.
A core thesis at Physical Intelligence is that developing a general-purpose robotic model will ultimately be easier than creating specialized models for narrow applications.
Sergey Levine believes that robotic intelligence should be developed in a general, body-agnostic way rather than being tied to a specific form factor like a humanoid.
Multimodal language models provide a viable path to imbuing robots with common sense by leveraging the vast knowledge they contain.
The core research challenge at Physical Intelligence is to combine the knowledge from generative AI with the ability of deep reinforcement learning to achieve superhuman performance.
Physical Intelligence uses a "chain of thought" process, where a robot verbally reasons about a task before acting, to unlock common sense knowledge from its pre-trained model.
The key data acquisition strategy for robotics is to develop systems that are useful enough to be deployed in the real world, enabling them to collect more data autonomously.
Physical Intelligence's models have surprisingly generalized to different robot embodiments, including those with multi-fingered hands, without requiring model changes or explicit prompting about the robot's form.
Physical Intelligence discovered that its models reached a point where they could be improved with high-level language-based supervision, indicating the performance bottleneck shifted from low-level motor skills to mid-level scene interpretation.
A major challenge in robotics is the lack of an internet-sized dataset for training models, unlike what is available for large language models.
Sergey Levine predicts that in the long run, robotics will enable medical and surgical applications with machines that are not limited to humanoid forms or even direct human control.
A major controversy in robotics is the dichotomy between approaches heavily reliant on simulation data, common for humanoid acrobatics, and those reliant on real-world data, common for manipulation tasks.