AION
Research paperComputer Vision · Multimodal Models1 source · Oct 8, 2026

S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.

Key points

  • Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views.
  • Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns.
  • We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment.
  • Experiments on the University-1652 and SUES-200 datasets demonstrate that S$^3$Geo consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.

Sources (1)

  • [1]S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:51 AM
    To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.
    Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views.

Extractive summary: sentences quoted from the sources.