ResearchResearch paperImage, Video & 3D Generation1 source · Oct 8, 2026

SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model

We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run.

Key points

  • A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory.
  • SepGen supports two complementary modes: generation and separation.
  • We evaluate generation on scenes synthesized from text, and separation on scenes rendered by other generators and on real recordings of speech, music, and sound effects.
  • Given an audio-mix and captions that carry the spoken lines, SepGen outperforms language-conditioned separators, most clearly on speech, and it keeps the lead when the lines are removed from the captions.

Sources (1)

  • [1]SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:51 AM
    We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run.
    A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory.

Extractive summary: sentences quoted from the sources.

Related