Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
DGX agentarXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter