Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization
We introduce Branching Latent Diffusion (BLD), which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction.
ProofPaper ↗
Key points
- Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding.
- Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence.
- BLD combines latent compression with branching token realization, where each latent is decoded by a local AR branch and all branches run in parallel.
- Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.
Sources (1)
- [1]Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token RealizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:08 AM
We introduce Branching Latent Diffusion (BLD), which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction.
Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding.
Extractive summary: sentences quoted from the sources.