Research
Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech
arXiv:2604.11417v1 Announce Type: cross Abstract: Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few in
arXiv:2604.11417v1 Announce Type: cross Abstract: Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that derives iconic gesture placement and intensity from text and emotion alone, requiring no audio input at inference time. The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression on the BEAT2 dataset, while remaining computationally compact and suitable for real-time deployment on embodied agents.
Related
- KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
- RAM: Recover Any 3D Human Motion in-the-Wild
- STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems
- Towards Green Wearable Computing: A Physics-Aware Spiking Neural Network for Energy-Efficient IMU-based Human Activity Recognition
Source: arXiv cs.AI | 2026-04-14