Reve 2.0 is betting on layouts to make images editable like code.

Reve 2.0 uses structured layouts to make 4K images editable like code, outperforming GPT-Image-1.5 to rank second in the Text-to-Image Arena.

Reve is releasing the second major version of its model. Rather than generating an image directly from a prompt, Reve 2.0 relies on an intermediate representation called Layout: a structured, hierarchical plan where each element receives a position, size, local description, and optional attributes like color or an image reference. The principle is reminiscent of what HTML is to a web page or SVG to a vector drawing: the image becomes readable and modifiable region by region, by a human as well as by an AI agent. Behind it, a unified Large Layout Model, built by fine-tuning open-source LLMs (Qwen) on billions of annotated images, accepts layouts, instructions, and reference images, then renders them in native 4K. Second in the Text-to-Image Arena, ahead of Nano Banana 2 and GPT-Image-1.5 and behind only GPT-Image-2, the model is claimed to be the world's best 4K generator, a level achieved, according to the team, with ten times fewer GPUs than major labs. Its internal ablations indicate that layout-based models outperform prompt-based generators of equivalent size, with quality improving as the number of described regions increases. The model can be tested at app.reve.com, with the API expected within a week.

User Feedback

The reception has been distinctly enthusiastic on 𝕏 as well as on Reddit, where discussions multiplied from the very first hours. What strikes testers is not so much the raw quality as the change in method: moving from a battle of prompts to designing a layout that the AI renders, a logic compared by some to AI-era DTP (Desktop Publishing) and perceived as a production tool rather than a random generator. Creatives particularly appreciate the ability to intervene: selecting a region, modifying it, and iterating without image degradation. Drawbacks remain isolated. A designer from the film industry places the model behind Grok Imagine and Midjourney for a 70s Giallo aesthetic, indicating that the result heavily depends on the requested style.