AION
Research paperImage, Video & 3D Generation · Computer Vision · Multimodal Models2 sources · Oct 7, 2026

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.

Key points

  • We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale.
  • We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space.
  • We fine-tune both families for monocular depth estimation and for image restoration/super-resolution.
  • We find no significant improvement from using a pixel-space generative prior.

Sources (2)

Extractive summary: sentences quoted from the sources.