
Qwen-Image-2.1: 7B open weights for generation, editing and alpha layers, under a research-only license
Qwen's Qwen-Image-2.1 folds text-to-image generation, image editing and native RGBA transparency into one 7B model whose weights are on Hugging Face and ModelScope, under a research license that bars commercial use.
Alibaba's Qwen team released Qwen-Image-2.1 on September 20: one image model that both generates and edits, and the first in the Qwen-Image line to handle transparent backgrounds natively. 1 The weights are on Hugging Face and ModelScope, with day-zero support in diffusers, ComfyUI, vLLM-Omni, SGLang and LightX2V. 23
The terms are where the release narrows. Qwen calls it open weights, but the weights ship under the Qwen Research License Agreement, which grants use "for research or evaluation purposes only" and says a commercial use needs a separate license. 4 On Qwen's own Qwen-Image-Bench the model scores 60.28 overall — seventh of the 24 systems the chart ranks, behind GPT Image 2.5 Sunburst at 67.01 and behind Qwen's own hosted Qwen Image 3 Pro at 62.36. 13
| Signal | Confirmed detail | What it means for you |
|---|---|---|
| What shipped | A 7B-parameter visual generation component with 32 single-stream DiT layers and native 2K output; ModelScope lists 16.22B parameters for the whole pipeline, the Qwen3-VL 8B text encoder included. 15 | One prosumer GPU runs it with CPU offloading; the download is 33 GB. |
| What's new | Native RGBA throughout: generate a transparent image from a prompt, edit text inside a transparent layer, or lift a subject out of an ordinary photograph as a cut-out layer. 1 | It absorbs December 2025's separate Qwen-Image-Layered model; compositing no longer needs a background-removal pass. |
| Editing | Up to 10 reference images; local edits marked with circles, painted strokes or a separate mask; portrait and product identity preserved. 16 | Try-on, interiors and storyboards come from one model, not a chain. |
| The catch | The research license is non-commercial, and redistribution must carry the agreement and a Notice file. 4 | Free to evaluate, not to ship; a product needs a commercial license first. |
What the model actually is
Qwen-Image-2.1 is a single-stream diffusion transformer with block-causal attention: the system prefix and the edit instruction get a token-level causal mask, images get a chunk-level one. 2 The mixed granularity lets the model encode its reference images and instruction once and reuse that cached prefix on every denoising step — the source of Qwen's efficiency claim, which pays off most on edits carrying several input images. 2 A 64-channel RGBA autoencoder at 16× compression carries the alpha channel, and two fine-tuned Qwen3.5-VL 9B checkpoints ship beside the model to expand short prompts. 2
Where it stands, and where it stops

Qwen's framing is a 7B model that "outperforms most closed-source models", which its own chart roughly supports for the systems ranked below it. 3 What the chart cannot settle is how the model does on a yardstick Qwen does not own, and no third-party evaluation exists yet: the release is hours old.
Before you build on it
Transparency is prompt-driven, not a switch: the model decides from the prompt whether to emit an ordinary image or one with an alpha channel, and the card says to ask for it outright. Everything else stays on default: 40 denoising steps, 2048×2048 output. 26
The license obligations outlive the download. Redistribution must carry the agreement and a Notice file; improving another model you distribute with these weights requires "Built with Qwen" in its documentation; and the agreement is governed by Chinese law, with disputes reserved to courts in Hangzhou. 4
References
- 1
- 2QwenLM/Qwen-Image-2.1
github.com
- 3Qwen on X
x.com
- 4Qwen Research License Agreement
huggingface.co
- 5Qwen-Image-2.1 on ModelScope
modelscope.cn
- 6Qwen/Qwen-Image-2.1
huggingface.co
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Grok 4.7 holds Grok 4.6's price and speed; independent testing puts it 7 points behind the leaders
- StepFun's Step 5 Preview matches Grok 4.6 and Kimi K3 on the intelligence index, at $0.71 a task
- Google's Gemini reached three real companies in May, from an evaluation environment that was supposed to be offline
- TypeSafe's Jev answers typed questions in 70–500 ms at $0.042 per million tokens, and ships with the vendor's own failure list
- Grok Voice Transcribe 2.0: a 2.7% word error rate at $0.20 an hour, while the API default still ships v1