Keep pulling the thread on Jim Fan.
The robotics field can follow the successful paradigm of large language models by simulating the next physical world state, aligning models through action fine-tuning, and using reinforcement learning for the final stage of training.
Video models like V03 can learn physical concepts such as gravity, buoyancy, lighting, reflection, and refraction emergently by predicting subsequent pixels at a large scale.
NVIDIA's DreamZero is a new type of policy model that can zero-shot solve tasks and verbs it has never seen in training by jointly decoding future world states and actions.
Jim Fan proposes that World Action Models (WAMs) should replace Vision Language Action (VLA) models as the dominant paradigm in robotics.
NVIDIA's Eagle Scale model was pre-trained on 21,000 hours of in-the-wild egocentric human video data with no robot data included.
The fine-tuning for NVIDIA's Eagle Scale model used only 50 hours of motion capture data and 4 hours of teleoperation data, which constituted less than 0.1% of the total training mix.
NVIDIA researchers have discovered a neural scaling law for dexterity, demonstrating a log-linear relationship between the hours of pre-training data and the model's optimal validation loss.
Jim Fan predicts that the volume of egocentric video data for robotics training will reach 10 million hours within the next year.
Jim Fan predicts that within the next one to two years, the use of teleoperation for robotics data collection will decrease to a negligible amount.
Jim Fan predicts that egocentric videos will become the primary data source for training robots, supplemented by custom-designed data wearables.
NVIDIA's DreamDojo is a neural simulator that generates interactive environments from video world models without using any classical physics equations or graphics engines.
Jim Fan predicts that robotics will pass the "Physical Turing Test," where a robot's task performance is indistinguishable from a human's, within the next two to three years.