Safety
CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment
arXiv:2608.21041v1 Announce Type: cross Abstract: Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite
arXiv:2608.21041v1 Announce Type: cross Abstract: Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel nderline{Co}ntrastive-based nderline{S}patial-nderline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7% over the strongest competing methods across eight city-indicator settings. The code is available in href{https://github.com/Arandinglv/CoST}{this repo}.
Related
- STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
- FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction
- RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment
Source: arXiv cs.AI | 2026-08-24