Research
Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
arXiv:2609.01136v1 Announce Type: new Abstract: Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images. Although different change t
arXiv:2609.01136v1 Announce Type: new Abstract: Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images. Although different change types (e.g., color shifts, object additions) exhibit distinct visual cues and require specialized reasoning processes, existing methods often overlook these distinctions. To address this limitation, we propose Multi-Expert Diagnosis for Image Change (MEDIC), a novel framework that introduces change-type awareness by explicitly modeling change categories. MEDIC employs type-specialized memory experts that dynamically retrieve type-relevant visual patterns conditioned on the input. This design enables each expert to capture diverse variations within its change type while focusing on the most informative visual cues. By softly routing inputs across type-specialized experts and learning dedicated representations for each change category, MEDIC generates more precise and type-aware change descriptions. Extensive experiments demonstrate that the proposed MEDIC consistently outperforms existing methods across diverse and challenging datasets. The code is available at href{https://github.com/VisualAIKHU/MEDIC}{GitHub}.
Related
- LBTCap: A Lightweight Bilateral Transformer for Real-Time Remote Sensing Image Change Captioning
- STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
- Freq-RemoteVAR: Next-Frequency Autoregressive Modeling for Remote Sensing Change Detection
- MoTE: Mixture of Task Experts for Multi-Task Video Understanding
Source: arXiv cs.CV | 2026-09-02