Research
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
arXiv:2608.13037v1 Announce Type: new Abstract: Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant ch
arXiv:2608.13037v1 Announce Type: new Abstract: Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing training-free methods produce solid overall results in spatially grounded generation, ie, placing a specific object in a designated location, they rely on gradient-based optimization techniques that incur prohibitive computational overhead, a bottleneck amplified in modern large-scale architectures. To address this limitation, we present Gradient-free Analytical Trajectory Optimization Video Generation (GATO-Vid), a novel training-free and gradient-free approach for precise spatial guidance. Rather than relying on costly backward passes, we introduce an alternative cross-attention score and solve it analytically to obtain an exact, closed-form solution. To use our analytical solution, we propose an on-the-fly injection mechanism tailored to the topological manifold of the transformer's latent space. Our experiments demonstrate that GATO-Vid significantly outperforms existing baselines in localization accuracy while introducing minimal computational overhead.
Related
- Compositional Video Generation via Inference-Time Guidance
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
- Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers
Source: arXiv cs.CV | 2026-08-14