Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images
We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training.
ProofPaper ↗
Key points
- Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning.
- While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models.
- The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference.
- We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models.
Sources (1)
- [1]Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide ImagesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:56 AM
We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training.
Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Oct 2, 2026Limits of Confidence in Diffusion
- Sep 30, 2026Cloudflare/clef-flash
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0