ResearchResearch paperMultimodal Models · Speech & Audio1 source · Oct 8, 2026

Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion

Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.

Key points

  • The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text.
  • Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN).
  • We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams.
  • While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained.

Sources (1)

  • [1]Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:34 PM
    Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training.
    The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text.

Extractive summary: sentences quoted from the sources.

Related