GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
DGX agentarXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan