Qwen-Robot is Here: Controls a Robot With Zero Training

Qwen-Robot is Here: Controls a Robot With Zero Training

More

Descriptions:

Alibaba’s Qwen team has released Qwen Robot Suite — three separate foundation models targeting the physical world — and this video from Fahd Mirza breaks down what each one does and why the architecture choices matter. RobotNav handles navigation, RobotManip handles manipulation, and RobotWorld is a video-based world model for predicting future states. All three are built on top of existing Qwen vision-language backbones (Qwen3-VL, Qwen3.5, Qwen2.5VL) with small action heads bolted on, meaning they share the same language-understanding foundation that developers have already been running locally.

The numbers behind the release are significant: RobotNav trained on 15.6 million samples across five task families, RobotManip pre-trained on over 38,000 hours of manipulation data (all from open datasets and human hand video), and RobotWorld built from 8.6 million video-text pairs spanning 20+ embodiment types. Demo footage shows a Unitree Go2 quadruped navigating an unseen apartment zero-shot at 196 milliseconds per step on a Jetson Thor chip, and a robot retracing a 22-meter route through an exhibition hall purely from language instructions.

One important caveat: the GitHub repo confirms no current plan to release weights for RobotNav or RobotManip, making this a papers-and-demos release for now. Mirza also teases a follow-up on RobotManip’s claim that standard robot manipulation benchmarks are fundamentally broken — a topic worth watching for anyone following embodied AI research.


📺 Source: Fahd Mirza · Published August 07, 2026
🏷️ Format: News Analysis

1 Item

Channels