Model Releases
Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming
arXiv:2608.25495v1 Announce Type: new Abstract: Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally
arXiv:2608.25495v1 Announce Type: new Abstract: Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human action recognition rely either on sparse skeletal representations, which lack fine-grained motion cues, or dense optical flow, which can be computationally expensive for low-latency perception pipelines. In this paper, we propose PoseOFF, a pose-anchored optical flow representation that captures local motion information around human joints to support earlier human intent understanding. By conditioning motion feature extraction on human pose, PoseOFF encodes localised motion dynamics at semantically meaningful body locations, forming a structured motion representation that is explicitly aligned with human kinematics. We evaluate PoseOFF across multiple benchmark datasets and backbone architectures for action anticipation, demonstrating consistent improvements in recognition accuracy, particularly at early observation ratios. Our results show that PoseOFF enables models to achieve comparable or improved performance while observing less of the action sequence, highlighting its effectiveness for early prediction. Importantly, these gains are achieved without requiring full-frame motion processing, making the approach practical for real-time and resource-constrained settings. These findings suggest that pose-centred motion representations such as PoseOFF can enhance the ability of interactive robot systems to infer human actions earlier, supporting more responsive and anticipatory behaviour in human-robot interaction scenarios.
Related
- HUI360: A 360{eg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation
- Zero-Shot Skeleton-Based Action Anticipation
- Action-Guided Attention for Video Action Anticipation
- H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
Source: arXiv cs.CV | 2026-08-27