BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image Diffusion
We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time.
ProofPaper ↗
Key points
- Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches.
- BudgetPix comprises three key components: (1) an adaptive encoder that maps a fixed-size image to a variable-length token sequence using an entropy-guided quadtree alongside a multi-scale patch embedder; (2) a scale-aware decoder reconstructs fixed-resolution images from multi-scale token sets; and (3) a flexible training and sampling schedule that enables pixel-space denoisers to operate across variable token counts.
- Evaluated on text-to-image generation, BudgetPix matches the fidelity of MiniT2I-L at $512^2$ and PixelDiT at $1024^2$ using just 25% of the original compute budget.
- In class-conditional generation using a MeanFlow backbone, BudgetPix requires merely 60% of the full compute budget to produce images with near-zero quality degradation, observing a marginal 0.8-point increase in FID.
Sources (1)
- [1]BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image DiffusionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:54 PM
We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time.
Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches.
Extractive summary: sentences quoted from the sources.