Model Releases

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

arXiv:2608.25367v1 Announce Type: new Abstract: Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, a

DGX agentpaper
model-releasesarxiv-cs-cv

arXiv:2608.25367v1 Announce Type: new Abstract: Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While, optical images have rich object structural information, and sonar images are less affected by underwater noise and have a longer visible distance. Optical (RGB modality) and sonar (Sonar modality) images have complementary information underwater. In this paper, we create an RGB-Sonar multimodal object detection dataset, extbf{R}GB-extbf{S}onar extbf{Fusion} (RSFusion) and propose evaluation metrics for the benchmark. And we propose the extbf{R}GB-extbf{S}onar extbf{Fusion} extbf{Det}ector (RSFusionDet) with a new RGB-Sonar multimodal object detection result expression for RGB-Sonar multimodal object detection. We analyze the features of RGB and Sonar modal information, and design a Cross-Attention Fusion (CAFusion) module to fuse RGB-Sonar spatial misalignment features and Object Matching Head (OMHead) with Loss (OMLoss) to match identical objects in RGB-Sonar modalities. Our RSFusionDet achieves 76.4/48.6 AP (RGB/Sonar) for object detection and 83.4 (ext{F1-Score}_{match}) for object matching, on RSFusion, which outperforms other object detection models. Compared with the DINO baseline, our method improves by 0.7/1.4 AP (RGB/Sonar) while simultaneously providing reliable cross-modal object matching. The code and datasets are publicly available at https://github.com/LEFTeyex/RSFusionDet.

Related

Source: arXiv cs.CV | 2026-08-27

Loading related sources…