QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation.
Key points
- Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility.
- The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution.
- Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation.
- Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities.
Sources (1)
- [1]QuadTok: Quadtree Visual Tokenizer for Autoregressive Image GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:47 PM
We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation.
Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility.
Extractive summary: sentences quoted from the sources.