AION
Research paperImage, Video & 3D Generation · Interpretability1 source · Oct 7, 2026

Pooling Representation Autoencoders for Efficient Diffusion

Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens.

Key points

  • Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive.
  • Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder.
  • On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency.
  • Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.

Sources (1)

  • [1]Pooling Representation Autoencoders for Efficient Diffusion
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:13 AM
    Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens.
    Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Consistent Distribution Matching for Data-Free Diffusion Distillation
  2. Oct 6, 2026From the Drosophila Visual Connectome to General-Purpose Computer Vision
  3. Oct 6, 2026Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration
  4. Oct 6, 2026Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
  5. Oct 6, 2026Later Is Better: Token Reduction for ViTs Under Distribution Shift
  6. Oct 1, 2026nvidia/PixelDiT2-ImageNet

Related