Research
Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts
arXiv:2608.13458v1 Announce Type: new Abstract: Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appear
arXiv:2608.13458v1 Announce Type: new Abstract: Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.
Related
- B-MoE: A Body-Part-Aware Mixture-of-Experts 'All Parts Matter' Approach to Micro-Action Recognition
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
- TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
- Geometry-Guided Self-Supervision for Ultra-Fine-Grained Recognition with Limited Data
Source: arXiv cs.CV | 2026-08-14