ResearchResearch paperEfficiency & Inference · Large Language Models1 source · Oct 7, 2026

Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization

We introduce Branching Latent Diffusion (BLD), which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction.

Key points

  • Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding.
  • Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence.
  • BLD combines latent compression with branching token realization, where each latent is decoded by a local AR branch and all branches run in parallel.
  • Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related