Model Releases
MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving
arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autono
arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks. However, such a single feature map hardly carries sufficient information to simultaneously meet the requirements of various perception tasks, leading to a very limited perception performance. To mitigate this limitation, this paper proposes MATS, a novel multi-modality multi-task learning approach with modality-adaptive BEV fusion and task-specific Mixture-of-Experts (MoE) for 3D perception. Specifically, a simple modality-adaptive BEV fusion module is designed to adaptively recalibrate the BEV features by modeling the global cross-modality dependencies, generating diverse BEV feature maps for various perception tasks. For joint multi-task learning, this paper proposes a task-specific MoE module to decouple the tasks and enable the network to automatically choose the appropriate BEV feature candidates for each specific task. To validate the effectiveness of the proposed approach, we conduct extensive experiments on the large-scale benchmark nuScenes. With the camera- and LiDAR-modality input data, the proposed approach outperforms the state-of-the-art (SOTA) by a significant margin. Furthermore, the experimental results on the single tasks show that the proposed approach significantly outperforms the baselines. The code and trained models will be available upon publication.
Related
- Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection
- FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog
- WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather
- CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks
Source: arXiv cs.CV | 2026-07-28