SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run.
ProofPaper ↗
Key points
- A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory.
- SepGen supports two complementary modes: generation and separation.
- We evaluate generation on scenes synthesized from text, and separation on scenes rendered by other generators and on real recordings of speech, music, and sound effects.
- Given an audio-mix and captions that carry the spoken lines, SepGen outperforms language-conditioned separators, most clearly on speech, and it keeps the lead when the lines are removed from the captions.
Sources (1)
- [1]SepGen: Multi-Stem Audio-Video Separation and Generation in a Single ModelarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:51 AM
We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run.
A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory.
Extractive summary: sentences quoted from the sources.