Model Releases
LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation
arXiv:2608.23880v1 Announce Type: new Abstract: Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating
arXiv:2608.23880v1 Announce Type: new Abstract: Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained with only image-level supervision that lacks guidance on which regions matter or how strongly each contributes. We propose LG-GER, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence, i.e., bounding boxes paired with emotion signals and confidence scores, for the training images. This structured evidence is distilled into a single vision-language model (VLM) backbone through four complementary losses: classification, region-text grounding, spatial emotion, and spatial confidence regression. At inference, LG-GER requires no detectors, no MLLM, and no multi-stream fusion, making GER practical for real-time and resource-constrained deployment. LG-GER has been evaluated on two benchmark GER datasets (GroupEmoW and GAF~3.0) and achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.
Related
- EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
- VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?
- Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning
Source: arXiv cs.CV | 2026-08-26