Keep pulling the thread on Sergey Levine.
Sergey Levin predicts that a robot capable of performing a useful, real-world task for people will be deployed within "single digit years," and hopes it will be within one or two years.
Modern LLMs and VLMs provide robots with a level of common sense that was unavailable five years ago, allowing them to make reasonable guesses about physical situations.
Sergey Levin provides a median estimate of five years for a robot to become capable of acting as a fully autonomous housekeeper.
Physical Intelligence's robotics model is a vision-language model adapted for motor control, featuring a vision encoder and an 'action expert' action decoder, structurally resembling a mixture-of-experts transformer.
Physical Intelligence's current robotics model operates with an inference speed of 100 milliseconds, a one-second context length, and a size of a few billion parameters.
Physical Intelligence's model generates continuous robot actions using a flow matching or diffusion-based method rather than representing actions as discrete tokens.
Robotics development in 2025 has a significant advantage over self-driving car development in 2009 due to vastly improved technology for generalizable and robust perception systems.
Sergey Levin speculates that future affordable robots may use a hybrid inference model, with a 'dumber reactive mode' running locally and more complex thinking offloaded to the cloud.
A Physical Intelligence robot demonstrated an emergent capability by picking up a second t-shirt that was obstructing its task and throwing it back in the bin, a behavior for which it was not explicitly trained.
Physical Intelligence is a company that aims to build general-purpose robotic foundation models capable of controlling any robot to perform any task.
Sergey Levin estimates that the amount of robotics data collected by Physical Intelligence is one to two orders of magnitude smaller than datasets used for training large multimodal models.
During its PIO5 project, Physical Intelligence discovered that its models could be effectively supervised with language instructions once they reached a certain level of competence.