Keep pulling the thread on Jim Fan.
The RoboTTT robot model and training recipe scales visuomotor context to 8,000 timesteps, which is three orders of magnitude beyond state-of-the-art policies.
The RoboTTT model's long context length enables new capabilities including one-shot in-context imitation from human video, on-the-fly policy improvement, and robustness to perturbations.
The RoboTTT research demonstrates for the first time that scaling pretraining context length leads to steady gains in closed-loop performance for robot policies.
On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over a single-step context baseline.
The RoboTTT model can fully complete a five-minute, ten-stage assembly task, which no baseline model could complete.
A RoboTTT model trained with an 8,000-timestep context outperforms the same model pretrained with a 1,000-timestep context by 62%.
The results from RoboTTT suggest that context length is a new scaling axis for robot foundation models.
The RoboTTT model scales visuomotor context without increasing inference latency.
RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies.
The recurrent state of the RoboTTT model consists of fast weights, which are parameters updated by gradient descent during both training and inference.