PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology
We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark.
ProofPaper ↗
Key points
- Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized.
- Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation component rather than a variable axis of model behavior.
- PathLang covers four task families: (1) zero-shot classification with image-text alignment analysis, (2) cross-modal retrieval, (3) paraphrase robustness, including semantic-equivalence paraphrases, length and reporting-style variation, and prompt ensembling, and (4) open-vocabulary diagnosis retrieval over four candidate pools with distinct forms of semantic competition.
- Across nine VLMs and five public datasets spanning four organs, we find that performance is highly sensitive to clinically equivalent paraphrases, varies substantially across forms of semantic competition, and that image-text alignment quality does not necessarily translate into inter-class separability.
Sources (1)
- [1]PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational PathologyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:27 AM
We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark.
Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized.
Extractive summary: sentences quoted from the sources.