Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
DGX agentarXiv:2607.08127v1 Announce Type: new Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these p