Model Releases
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
arXiv:2608.21160v1 Announce Type: new Abstract: Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human image
arXiv:2608.21160v1 Announce Type: new Abstract: Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.
Related
- MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts
- Do Visual Grounding Decoders Need Feed-Forward Networks? A Controlled Study over Frozen Vision-Language Features
- EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
- Open-Vocabulary Gaze Object Prediction: Benchmark and Method
Source: arXiv cs.CV | 2026-08-24