FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification
To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy.
Key points
- Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains.
- Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage.
- With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance.
- The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed.
Sources (1)
- [1]FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot ClassificationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:00 PM
To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy.
Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains.
Extractive summary: sentences quoted from the sources.