MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval
To address these issues, we propose a Multi-head Spiking Graph Attention Network (MSGAT) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation.
ProofPaper ↗
Key points
- Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning.
- However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment.
- However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations.
- Therefore, we introduce Sim-Fuse, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations.
Sources (1)
- [1]MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text RetrievalarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:56 AM
To address these issues, we propose a Multi-head Spiking Graph Attention Network (MSGAT) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation.
Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning.
Extractive summary: sentences quoted from the sources.