Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
Key points
- We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale.
- We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space.
- We fine-tune both families for monocular depth estimation and for image restoration/super-resolution.
- We find no significant improvement from using a pixel-space generative prior.
Sources (2)
- [1]Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-TuningHugging Face Daily Papers · Oct 7, 12:00 AM
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters.
We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a 256to512to1024 curriculum, after first ablating the prediction target and representation alignment at 256^2 to decide what to scale.
- [2]Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-TuningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:07 AM · same content
Extractive summary: sentences quoted from the sources.