AION
Research paperComputer Vision1 source · Oct 7, 2026

FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy.

Key points

  • Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains.
  • Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage.
  • With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance.
  • The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed.

Sources (1)

  • [1]FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:00 PM
    To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy.
    Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains.

Extractive summary: sentences quoted from the sources.