Pooling Representation Autoencoders for Efficient Diffusion
Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens.
Key points
- Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive.
- Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder.
- On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency.
- Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.
Sources (1)
- [1]Pooling Representation Autoencoders for Efficient DiffusionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:13 AM
Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens.
Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Consistent Distribution Matching for Data-Free Diffusion Distillation
- Oct 6, 2026From the Drosophila Visual Connectome to General-Purpose Computer Vision
- Oct 6, 2026Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration
- Oct 6, 2026Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
- Oct 6, 2026Later Is Better: Token Reduction for ViTs Under Distribution Shift
- Oct 1, 2026nvidia/PixelDiT2-ImageNet