HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions.
ProofPaper ↗
Key points
- Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions.
- Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation.
- This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations.
- HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture.
Sources (1)
- [1]HRIL: Learning Multimodal Synergy via Higher-Order Tensor ModelingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:39 PM
Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions.
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions.
Extractive summary: sentences quoted from the sources.