Qwen-Image-3.0 pushes image generation toward truly actionable content

The Qwen-Image-3.0 model generates dense infographics and readable text at 10 pixels. What this precision changes for document creation.

With Qwen-Image-3.0, Alibaba focuses its advancements on three key areas: content density, detail precision, and the breadth of knowledge utilized.

The model accepts prompts of up to 4,500 tokens. This allows it to produce complex compositions in a single generation, such as an infographic grid, a storyboard, a newspaper page, educational material, or multiple nested interfaces. In its demonstration, Qwen notably generates nine distinct technical contents within a single image, complete with text, formulas, diagrams, and illustrations.

Precision also improves on small elements. The team claims readable text down to 10 px, including in scientific articles containing equations, subscripts, and special characters. The model also works on fine textures, such as skin, hair, paper, or brushstrokes, and can restore damaged artwork while preserving its original style.

Qwen-Image-3.0 natively supports twelve languages, over one hundred artistic styles, and various types of interfaces, from websites and video games to livestreams. It can also enrich an existing image with scientific annotations, taxonomic information, or elements from an online search.

Generation and editing are accessible from Qwen Studio.