AION
Research paperEfficiency & Inference · Image, Video & 3D Generation · Computer Vision1 source · Oct 8, 2026

CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding

Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces.

Key points

  • We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor.
  • CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed.
  • Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%.
  • On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities.

Sources (1)

  • [1]CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:07 AM
    Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces.
    We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor.

Extractive summary: sentences quoted from the sources.