Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
Key points
- Conventionally, the streaming video literature obtains both capabilities from a chained pipeline.
- Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel.
- As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training.
- The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality.
Sources (1)
- [1]Real-Time Joint Audio-Video Generation by Parallel Adapter CompositionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:23 PM
Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
Conventionally, the streaming video literature obtains both capabilities from a chained pipeline.
Extractive summary: sentences quoted from the sources.