ResearchResearch paperMultimodal Models · Speech & Audio · Interpretability1 source · Oct 8, 2026

Learning What to Trust in Multimodal Learning under Noisy Supervision

Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection.

Key points

  • Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions.
  • While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully.
  • This motivates us to build a more reliable noise detector in multimodal learning.
  • Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class.

Sources (1)

  • [1]Learning What to Trust in Multimodal Learning under Noisy Supervision
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:14 AM
    Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection.
    Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions.

Extractive summary: sentences quoted from the sources.

Related