Keep pulling the thread on Junyang Lin.
Qwen-RobotNav is robust to any inference-time configuration and does not require architectural modifications due to its training-time randomization over all parameters.
Qwen-RobotNav's parameterized interface allows an upper-level planner to dynamically switch its task mode and context strategy mid-episode for long-horizon scenarios.
Qwen-RobotNav achieves new state-of-the-art results across major navigation benchmarks.
Qwen-RobotNav demonstrates strong zero-shot generalization to real-world robots across diverse environments.
The Qwen-RobotNav navigation model has a parameterized interface with two dimensions: multiple task modes for selecting behavior and controllable observation parameters like token budget and per-camera weights for managing visual history.
The Qwen-RobotNav model was trained on a dataset of 15.6 million samples.
Co-training Qwen-RobotNav with vision-language data prevents the model from collapsing into a reactive action-sequence mapper, a problem observed in trajectory-only training.
The Qwen-RobotNav model demonstrates favorable scaling properties when its parameter count is increased from 2 billion to 8 billion.
The joint multi-task training of Qwen-RobotNav develops a shared spatial-planning substrate that transfers across different task families.