Prior Directions: Why GUI Grounding Gets Locked in the Past
DGX agentarXiv:2607.26913v1 Announce Type: new Abstract: Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language