Local Ai
XRF-to-Optical Field-of-View Localization with Vision Language Models
arXiv:2608.18309v1 Announce Type: new Abstract: Registering images acquired with different microscopy modalities is essential for relating complementary measurements of the same specimen. In correlati
arXiv:2608.18309v1 Announce Type: new Abstract: Registering images acquired with different microscopy modalities is essential for relating complementary measurements of the same specimen. In correlative X-ray fluorescence (XRF) and optical microscopy, the XRF map often covers only a small region of an optical image acquired from the same or an adjacent tissue section. Field-of-view (FOV) localization is necessary but can be difficult when appearance and structure differ across modalities. Here we evaluate training-free vision language model (VLM) localization on two datasets representing same-section high-correspondence and adjacent-section low-correspondence imaging. We test unconstrained and metadata-constrained search and compare VLMs with geometric controls, classical template matching, and two alternative training-free approaches (DINOv2 and multiGradICON). Direct VLM prompting produced content-dependent spatial signals but was not reliable alone. Classical matching was most accurate when cross-modal structure was preserved but failed in the low-correspondence collection. A proposal-and-verify workflow used repeated VLM predictions as candidates and image-based similarity to select the final location. This workflow recovered useful localization in the low-correspondence regime.
Related
- More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization
- Mechanisms of Object Localization in Vision-Language Models
- Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
Source: arXiv cs.CV | 2026-08-20