Iris-3B wants to generate the image directly, without going through a latent space
Speridlabs' Iris-3B model generates images directly in pixel space. A 3-billion-parameter open-weight alternative to latent models.
Speridlabs presents Iris-3B, a 3-billion-parameter open-weight text-to-image model that takes a technical direction still relatively rare at this scale: generating the image directly in pixel space, without using a VAE responsible for compressing and then reconstructing the visual. The project is entirely open, with weights available on Hugging Face, code, a demo, and an Apache 2.0 license.
The proposition is interesting because it challenges one of the building blocks that has become almost standard in modern image generation. Latent diffusion models like Stable Diffusion or several more recent architectures generally do not work directly on the final image. They first compress this image into a more compact representation, perform the generation in this space, and then use a decoder to return to the visible rendering.
Iris-3B attempts to eliminate this intermediate step. Why Remove the VAE?
The main advantage of latent models is their efficiency. Working on a compressed representation greatly reduces the amount of computation required to generate a high-resolution image.
However, this compression comes with a trade-off: some fine information can be lost or distorted even before the generative model intervenes. Speridlabs points out textures, small details, and certain precise structures in particular, which can be affected by the biases introduced by the latent representation.
The paper dedicated to Iris-3B therefore explores the hypothesis that a model trained directly in pixel space could better preserve this information. The authors do not limit themselves to text-to-image generation: they also test this representation on depth estimation and image restoration or super-resolution.
However, this hypothesis does not lead to a clear victory. The researchers themselves indicate that they do not observe a significant and systematic improvement simply from removing the latent space on these downstream tasks. On depth estimation, Iris-3B achieves a level comparable to a latent variant of FLUX.2 Klein, while on DIV2K ×4 restoration, the variants working directly in pixel space do not outperform the corresponding latent model.
This is probably one of the most interesting aspects of the project: Speridlabs also publishes the results that do not confirm its initial hypothesis. A 3-Billion-Parameter Transformer
Iris-3B is a 3-billion-parameter flow-matching transformer.
Images are divided into 16 × 16 patches, then processed by an architecture composed of eight dual-stream blocks, where text and image have separate weights, followed by sixteen single-stream blocks. A small Pixel Transformer head, or PiT head, then directly reconstructs the values of each patch into image form.
The text is encoded by a frozen Qwen3-VL-4B-Instruct, used solely as an encoder. The image model itself therefore has 3 billion parameters, to which this external text encoder is added.
Speridlabs also uses a 2D RoPE positional representation, which allows the same set of weights to work with different aspect ratios and resolutions up to approximately one megapixel.
The model was trained from scratch using a progressive curriculum: 256 × 256, then 512 × 512, then 1024 × 1024, before a supervised fine-tuning phase at 1024. The Hugging Face model card indicates a total of 665,000 training steps. Images Generated Directly in 1024
In standard use, Iris-3B produces images around 1024 × 1024, with about a hundred generation steps and a CFG set to 3 in the default configuration. The weights represent about 12 GB and currently require an NVIDIA GPU with CUDA for the local execution proposed by Speridlabs.
The model also accepts native aspect ratios up to approximately one megapixel rather than systematically imposing a square.
As for prompting, there is nothing particularly exotic: Speridlabs recommends simple descriptions in English and supports negative prompts, fixed seeds, and batch processing.
However, the company acknowledges several limitations familiar to current image models. Rendering text within the image is not always reliable, nor is strictly respecting the requested number of objects. For large prints, it still recommends using an upscaling tool. Performance Close to Much Larger Models
On benchmarks published by Speridlabs, Iris-3B remains competitive against several heavier latent architectures.
The model achieves 0.798 on GenEval, 86.52 on DPG, 0.857 on LongText, and 0.540 on OneIG.
On OneIG, it is practically on par with Qwen-Image 20B, reported at 0.539 in the table used by Speridlabs, even though Iris-3B has about seven times fewer parameters in its generative backbone. However, it remains behind Qwen-Image on GenEval, DPG, and LongText.
Z-Image 6B also obtains better scores on several of these metrics, but Speridlabs specifies that its own results are obtained without prompt rewriting and without a best-of-N strategy.
These comparisons should be treated with caution, as some models use different protocols and not all results necessarily come from a perfectly homogeneous evaluation campaign. Generating and Understanding with the Same Backbone
The other goal of Iris-3B is to test whether a generative model can also serve as a foundation for visual understanding tasks.
Speridlabs thus presents it as a general vision learner and compares this approach to that of specialized models like DINOv2.
The model is adapted for monocular depth estimation as well as image restoration and super-resolution. The idea is that the representation acquired by learning to directly reconstruct every detail of an image could be reused to analyze that same visual structure.
However, the paper's results temper this intuition. Switching to pixel space is not enough on its own to improve depth or restoration tasks compared to properly tuned latent architectures.
This does not invalidate the approach, but it shows that removing the VAE is not an automatic solution. The Return of Pixel-Space at Scale
The approach of Iris-3B is part of a broader resurgence of