Safety
SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception
arXiv:2608.10497v1 Announce Type: new Abstract: While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature ex
arXiv:2608.10497v1 Announce Type: new Abstract: While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from "semantic blindness," overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.
Related
- Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
- Identifying Ethical Biases in Action Recognition Models
- Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Models
- Improving Reasoning in Vision-Language Models via Perception Verified Self-Training
Source: arXiv cs.CV | 2026-08-12