Qwen-Image-2.1
View on GitHubQwen's most powerful open-source image generation model
Qwen's 7B text-to-image and image-editing diffusion model (32-layer single-stream DiT) with native RGBA transparency, up to 10 reference images, mask/local edits, and 2K output. Ships Diffusers, ComfyUI, vLLM-Omni and SGLang integrations plus prompt-rewriting checkpoints.
Use Cases
Text-to-image generation at native 2K resolutionInstruction-based image editing with masks or painted regionsTransparent/RGBA sticker and layer generationMulti-subject composition from up to 10 reference imagesSubject extraction from photographsShort prompt rewriting into detailed generation promptsHigh-throughput batch image generation with vLLM offline inferenceOnline image generation serving via OpenAI-style HTTP APIIdentity-preserving edits for people and products
Built With
- Language
- Python
- Frameworks
- PyTorch · HuggingFace Diffusers · Transformers · vLLM · vLLM-Omni · SGLang · ComfyUI · LightX2V
Tags
text-to-image · image-editing · diffusion · DiT · RGBA-transparency · 7B-parameters · prompt-rewriting · 2K-resolution · multi-reference · FP8-quantization · ComfyUI · Diffusers · vLLM · SGLang · subject-extraction · inference-optimization