Back
Qwen-Image-2.1: a compact 7B model unifies image generation and editing
SiTech Team3 წთ. საკითხავი

Qwen-Image-2.1: a compact 7B model unifies image generation and editing

Qwen has open-sourced Qwen-Image-2.1, an image model that merges text-to-image generation and editing in a single network, with 7B parameters in its visual generation component and native transparency support.

A compact 7B architecture

Qwen has open-sourced Qwen-Image-2.1, an image model the team describes as balancing generation quality, inference efficiency and cost. It unifies text-to-image generation and image editing in one model whose visual generation component holds 7B parameters, with native support for transparent images.

That component is built from 32 Single-Stream DiT layers. Qwen says the compact size does not cost generation quality, and that inference efficiency was a focus for editing with multiple input images. The gains come from a mixed-granularity attention architecture: text, including the system prefix and editing instructions, uses a token-level causal mask, while image generation uses a chunk-level mask. KV cache reuse treats input images and editing instructions as static context, computed and cached in the first step, which speeds up inference and lowers memory use.

Native transparency in one model

Qwen-Image-Layered, introduced in December 2025, was a dedicated model for transparent image generation. Qwen-Image-2.1 folds that ability into the unified model and lets the prompt decide whether the output is a regular image or one with a transparency channel. Because generation and editing are unified, transparent images can be edited directly — for example, changing a subject's expression while keeping the background transparent. The same applies to photographs: from an RGB image the model can extract the chosen subject as an RGBA layer, making elements easier to reuse in later design and composition work.

Editing with up to 10 reference images

Qwen lists four editing improvements: multiple references, local edits, fidelity preservation and task coverage. The model accepts up to 10 reference images and can combine separate subjects and assets into one composition — six portraits merged into a group photograph, an outfit assembled from five references, a room arrangement built from 10 images of furnishings.

Local edits can be specified with circles, painted annotations or a separate mask; to leave the input intact, the original image and a mask can also be passed as two inputs. Fidelity work targets portrait identity and product consistency, keeping facial features stable across edits while preserving a product's text, textures and shape. Task coverage spans panoramas generated from a selfie, infographics expanded from a model photograph, and storyboards built from a three-view character reference.

The release also refines typography and portrait rendering; text rendering takes type styles and layout into account. Qwen-Image-2.1 is available on GitHub, Hugging Face and ModelScope, and the team presents it as a practical tool for design, content creation, e-commerce and visual storytelling.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.