AION
Research paperImage, Video & 3D Generation1 source · Oct 7, 2026

Real-Time Joint Audio-Video Generation by Parallel Adapter Composition

Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.

Key points

  • Conventionally, the streaming video literature obtains both capabilities from a chained pipeline.
  • Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel.
  • As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training.
  • The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality.

Sources (1)

  • [1]Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:23 PM
    Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
    Conventionally, the streaming video literature obtains both capabilities from a chained pipeline.

Extractive summary: sentences quoted from the sources.