Safety

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision

DGX agentpaper
safetyarxiv-cs-cv

arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision-language models such as CLIP have demonstrated remarkable semantic alignment capabilities, their application to continuous regression tasks is often limited by the discrete nature of text prompts. In this paper, we propose a novel multimodal framework for VA estimation that introduces Distance-aware Soft Prompt Guidance to bridge the gap between semantic representations and continuous affective dimensions. Specifically, we partition the VA space into multiple discrete regions, each associated with distinct textual descriptions. Rather than relying on hard categorization, we employ a Gaussian kernel to compute soft labels based on the Euclidean distance between the ground-truth coordinates and the region centers, allowing the model to learn fine-grained emotional transitions. For multimodal integration, our architecture utilizes a CLIP image encoder and an Audio Spectrogram Transformer to extract robust visual and acoustic features. These features are temporally modeled using Gated Recurrent Units and integrated through a hierarchical fusion scheme that sequentially combines cross-modal attention for alignment and gated fusion for adaptive refinement. Experimental results on the Aff-Wild2 dataset show that the proposed semantic-guided approach outperforms the official baseline and demonstrates robust performance on in-the-wild data.

Source: arXiv cs.CV | 2026-08-03

Loading related sources…