Keep pulling the thread on Junyang Lin.
The pretraining corpus for Qwen-RobotManip consists of approximately 38,100 hours of data.
Qwen-RobotManip demonstrates emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer.
Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $π$0.5, across all tested out-of-distribution (OOD) settings.
Qwen-RobotManip ranked 1st in the RoboChallenge benchmark with a 20% relative performance improvement.
The Qwen-RobotManip model has been validated on the AgileX ALOHA, Franka, UR, and ARX real-world robot platforms.
The Qwen-RobotManip model is built on the Qwen-VL foundation model.
Qwen-RobotManip uses a unified alignment framework across representation, motion, and behavioral dimensions to enable coherent, large-scale training from multiple data sources.
A human-to-robot synthesis pipeline was developed to convert egocentric hand demonstrations into robot trajectories for 15 different robot platforms.
The pretraining data for Qwen-RobotManip was sourced exclusively from open-source datasets and human videos, without using any proprietary data.
The performance of Qwen-RobotManip was evaluated on out-of-distribution (OOD) benchmarks, including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE.