GPT-Image 2.0
OpenAI launches GPT-Image 2.0 in ChatGPT, featuring improved text rendering, photorealism, and a pre-generation reasoning mode for 4K image production.
OpenAI launches GPT-Image 2.0, its new image generation and editing model integrated directly into ChatGPT. The main breakthrough concerns text rendering within images, which is now almost reliable, including for paragraphs, interfaces, or multilingual content. The model also significantly improves photorealism (skin, lighting, materials), adherence to complex instructions, and image-to-image editing, with the ability to modify elements while maintaining visual consistency and lighting. GPT-Image 2.0 also introduces a pre-generation reasoning mode, allowing it to analyze the request, plan the composition, and iterate before producing the final image. This logic brings the tool closer to a production system rather than a simple generator. The model supports outputs up to 4K and integrates via API (gpt-image-2) as well as on third-party platforms. The community highlights a clear gain in productivity, despite an aesthetic sometimes deemed more “clean” and slightly less expressive. I conducted my first tests yesterday late afternoon, and I will continue over the next few hours, but here's a recap: The model delivers on its promises regarding consistency (characters, poses, composition) and text adherence, marking a significant step forward for mockups and editorial uses. When tested with retail requests, fidelity remains partial for complex products. In cosmetics, visual artifacts (stray lines) persist in some generations, particularly on models' faces. The overall rendering still lacks sharpness in direct output, with a sense of constrained quality; this needs to be checked on the API side. As it stands, it performs well for structuring and iterating, while preserving the usual post-production workflow to achieve an exploitable final level. See my initial tests here.