Research
Pointing-Based Object Recognition
arXiv:2603.15403v2 Announce Type: replace Abstract: This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interactio
arXiv:2603.15403v2 Announce Type: replace Abstract: This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interaction moves toward more intuitive interfaces, the ability to identify targets of non-verbal communication becomes crucial. Our proposed system integrates several existing state-of-the-art methods, including object detection, body pose estimation, monocular depth estimation, and vision-language models. We evaluate the impact of 3D spatial information reconstructed from a single image and the utility of image captioning models in correcting classification errors. Experimental results on a custom dataset show that incorporating depth information significantly improves target identification, especially in complex scenes with overlapping objects. The modularity of the approach allows for deployment in environments where specialized depth sensors are unavailable.
Related
- Skarimva: Skeleton-based Action Recognition is a Multi-view Application
- Proper Body Landmark Subset Enables More Accurate and 5X Faster Recognition of Isolated Signs in LIBRAS
- Zero-shot Human Pose Estimation using Diffusion-based Inverse solvers
- A Multimodal Depth-Aware Method For Embodied Reference Understanding
Source: arXiv cs.CV | 2026-07-23