S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization
To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.
Key points
- Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views.
- Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns.
- We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment.
- Experiments on the University-1652 and SUES-200 datasets demonstrate that S$^3$Geo consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.
Sources (1)
- [1]S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-LocalizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:51 AM
To address these challenges, we propose S$^3$Geo, a structure-semantic synergistic learning framework for cross-view matching.
Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views.
Extractive summary: sentences quoted from the sources.