RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models
To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification.
ProofPaper ↗
Key points
- Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis.
- However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories.
- Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions.
- Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods.
Sources (1)
- [1]RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:28 PM
To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification.
Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 22, 2026vllm-project/vllm v0.30.0
- Sep 9, 2026vllm-project/vllm v0.29.0
- Aug 26, 2026vllm-project/vllm v0.28.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0