Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
ProofPaper ↗
Key points
- The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text.
- Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN).
- We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams.
- While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained.
Sources (1)
- [1]Open-Vocabulary Audio-Visual Event Localization via Complex-Valued FusionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:34 PM
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text.
Extractive summary: sentences quoted from the sources.