Alibaba's Qwen team released Qwen-Drive-1.0-4B, an open-source vision-language model for autonomous driving, in early September. Developed together with Huazhong University of Science and Technology, it folds three jobs into a single model: 3D perception of the surroundings, question answering about driving scenes, and generation of the vehicle's future path. Code, weights and demo data ship under the Apache 2.0 license.

Three Jobs on One Foundation

Qwen-Drive-1.0 is built on Qwen3.5-4B, the team's natively multimodal vision-language model. That architecture is left untouched, and two external modules are bolted on instead.

The first is the BEV Perception Head. Working in BEV (Bird's Eye View, a top-down plan of the area around the car), it handles 3D object detection, semantic occupancy prediction and map segmentation together. It acts as a probe for how much 3D information can be pulled out of the shared representations, and at the same time as an inspectable window onto the 3D structure of the scene.

The second is the Planning Expert, which takes the same representations and outputs the trajectory the ego vehicle is about to follow. The output is a sequence of 50 points covering 5 seconds at 10 Hz (10 samples per second), each carrying an x coordinate, a y coordinate and a heading.

Because the original language decoder is left in place, the model still answers both driving questions and general image questions. Training mixes perception, language and planning objectives in stages, and the team says it also standardized trajectories from several public driving datasets into a shared representation as part of the data pipeline work.

The Reinforcement-Learned Planner Comes Out Ahead

Two versions of the Planning Expert are distributed: planner-sft, trained to imitate driving examples, and planner-rl, further optimized against rewards.

On planning metrics, PDMS on NAVSIM v1.1 navtest is 88.2 for sft and 90.7 for rl. Under a best-of-6 protocol those rise to 89.3 and 91.4. RFS on the Waymo Open Dataset end-to-end test comes in at 7.78 and 7.91, again with rl on top.

The open-loop evaluation on NVIDIA PhysicalAI flips the order, though: minADE at 3 seconds is 0.34 m for sft against 0.38 m for rl. Which version you should pick depends on which metric you weight most heavily. Note that planner-rl was reward-optimized only on reasoning-conditioned rollouts, so it needs to be run in that mode. planner-sft is the one that covers both modes.

Learning to Drive Without Losing General Vision

The interesting part is how little of the base model's general ability was sacrificed for driving.

On driving-scene question answering, the plain Qwen3.5-4B scores 70.4 on LingoQA, while Qwen-Drive-1.0-SFT reaches 77.8. Ego3D RMSE, the error in 3D position estimation, shrinks from 13.17 to 7.78, and the causal-chain CoC score jumps from 2.6 to 41.3, a change of an order of magnitude.

Meanwhile general image understanding barely moves: MMStar goes from 75.3 to 75.9 and RealWorldQA from 76.3 to 79.0, with MMMU slipping only from 73.4 to 72.7. The familiar side effect where driving-specific fine-tuning wrecks general visual understanding has largely been avoided.

What It Takes to Run Locally

Everything ships in one directory: 9.1 GB for the vision-language model itself, 2.1 GB each for planner-sft and planner-rl, and 0.5 GB for the perception head. The head you need is specified and attached when the model is loaded.

A GPU with 24 GB or more of memory is recommended. The repository bundles four demo scenes — a night intersection whose light turns green, a left turn, a right turn, and slowing down past a parked truck — so the behavior can be checked with nothing but the weights.

export PYTHONPATH=src
python scripts/demo.py --model Qwen-Drive-1.0-4B --planner Qwen-Drive-1.0-4B/planner-rl \
    --scenes data/demo/planning_scenes.jsonl --image-archive data/demo/frames.parquet \
    --plot demo.png

Running this command writes out a single figure combining the camera views of the scene, the predicted trajectories against the ground truth, and the generated reasoning. The weights are available from either Hugging Face or ModelScope.

Summary

Qwen-Drive-1.0-4B is an open-source driving model that adds 3D perception and trajectory generation on top of a general-purpose vision-language model. The reinforcement-learned planner leads on most planning metrics, though the imitation-trained one wins on others, and general image understanding is essentially preserved. With the full set of weights under Apache 2.0 and a 24 GB-class GPU enough to reproduce the behavior locally, it looks like an easy foundation to build research and evaluation work on.