What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
DGX agentarXiv:2605.20795v1 Announce Type: new Abstract: Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-bas