Research
Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis
arXiv:2607.12686v1 Announce Type: new Abstract: Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled
arXiv:2607.12686v1 Announce Type: new Abstract: Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose extbf{SeRIn} (extbf{Se}gregate, extbf{R}efine, extbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Related
- MEME-Fusion@CHiPSAL 2026: Multimodal Ablation Study of Hate Detection and Sentiment Analysis on Nepali Memes
- Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
- Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis
Source: arXiv cs.CL | 2026-07-15