【#Tech24H】On July 21, Alibaba released its latest foundational image generation model: Qwen-Image-3.0. As the third-generation foundational model in the Qwen series for image generation and editing, it features comprehensive capability upgrades. Scenes such as livestream room pages that GPT-Image-2 excelled at can now also be generated by Qwen-Image-3.0. The new model supports ultra-long inputs of up to 4.5k tokens, enabling the one-shot generation of knowledge diagrams that integrate formulas, symbols, geometric figures, logical derivation steps, and other multiple elements, as well as complex UI interfaces. It also natively supports rendering in 12 languages and over 20 font families, reducing production costs for “readable and usable” assets such as multilingual product posters, film and television storyboards, and comic panels. In addition, Qwen-Image-3.0 possesses richer knowledge and can readily comprehend complex instructions and nested logical relationships within scenes. Following a single instruction, the model can generate a composite visual expression within one image that simultaneously incorporates web pages, software interfaces, chat windows, posters, and other diverse visual content. [ By Zhang Liyan | Tang Ruohan ]

