ResearchResearch paperComputer Vision1 source · Oct 6, 2026

RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models

To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification.

Key points

  • Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis.
  • However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories.
  • Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions.
  • Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 22, 2026vllm-project/vllm v0.30.0
  2. Sep 9, 2026vllm-project/vllm v0.29.0
  3. Aug 26, 2026vllm-project/vllm v0.28.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jul 11, 2026vllm-project/vllm v0.25.0
  6. Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0

Related