Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
arXiv:2607.18789v1 Announce Type: new Abstract: Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture