The Qwen team, the generative AI brand under China's Alibaba, has released a new image generation and editing model called Qwen-Image-2.1 under a free research license. Despite a lightweight 7B parameter design, the model supports generating and editing images with transparent backgrounds, along with editing based on up to 10 reference images at once.

A 7B Design That Handles Both Generation and Editing

Qwen-Image-2.1 combines a 7B-parameter image generation component with an 8-billion-parameter vision-language model, Qwen3-VL, which handles image and text understanding. Internally, it uses a 32-layer diffusion transformer (DiT) architecture that treats text and image information as a single sequence, along with a mechanism that adjusts its attention range depending on the task, making it easier to follow complex instructions. For text-to-image generation, the model outputs images at a native resolution of 2048x2048 pixels, supporting seven aspect ratios up to 2752x1536 pixels. Its editing feature accepts up to 10 reference images, allowing users to change a subject's pose or background while preserving identity, or to rewrite only a specified region of an image. The model can also generate and edit RGBA images with native transparent backgrounds, and it reuses cached computations from text and reference images to improve processing efficiency.

Alibaba Claims Its Benchmark Beats Google's "Nano Banana 2.0"

On Alibaba's own evaluation suite, Qwen-Image-Bench, Qwen-Image-2.1 scored 60.28 points, edging out Google's image generation model Nano Banana 2.0, which scored 59.82. Alibaba says this places the model at the top among open-weight models currently available, claiming performance comparable to image generation models from OpenAI and Meta. That said, the figures come from a benchmark Alibaba designed itself, and independent third-party verification has not yet been widely conducted.

Commercial Use Requires a Separate License, Broad Tool Support

Qwen-Image-2.1 is released under the Qwen Research License, which allows free use for research and evaluation but requires a separate commercial agreement with Alibaba for commercial use. Rights to images generated with the model reportedly belong to the user and fall outside the scope of the license. The model can be downloaded from Hugging Face, ModelScope, and GitHub, and is already supported by major tools including Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V. Running it locally generally requires a GPU with at least 24GB of VRAM, and beyond NVIDIA hardware, it also runs on AMD's ROCm environment and several other chip platforms through FlagOS. The model supports multilingual prompts, including Japanese, and is already being tried out by developers and creators in Japan.

Summary

In September 2026, Alibaba released Qwen-Image-2.1, an image generation and editing model, under a free research license. Despite its lightweight 7B design, the model supports generating transparent backgrounds and editing with multiple reference images, and Alibaba claims its own benchmark places it ahead of larger rivals' models. While commercial use requires a separate license, the fact that it can be tried locally as an open-weight model is likely to broaden the options for companies and individuals looking to build image generation AI into their own products.