Model Releases

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

arXiv:2608.15456v1 Announce Type: new Abstract: Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most rema

DGX agentpaper
model-releasesarxiv-cs-cv

arXiv:2608.15456v1 Announce Type: new Abstract: Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.

Source: arXiv cs.CV | 2026-08-18

Loading related sources…