Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
arXiv:2605.15342v1 Announce Type: new Abstract: Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation