Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
arXiv:2607.04311v1 Announce Type: new Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity