Local Ai
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
arXiv:2604.20470v1 Announce Type: new Abstract: Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing
arXiv:2604.20470v1 Announce Type: new Abstract: Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long-range information in complex dynamics. To address this issue, we propose extbf{DynamicRad}, a unified sparse-attention paradigm that grounds adaptive selection within a radial locality prior. DynamicRad introduces a extbf{dual-mode} strategy: extit{static-ratio} for speed-optimized execution and extit{dynamic-threshold} for quality-first filtering. To ensure robustness without online search overhead, we integrate an offline Bayesian Optimization (BO) pipeline coupled with a extbf{semantic motion router}. This lightweight projection module maps prompt embeddings to optimal sparsity regimes with extbf{minimal runtime overhead}. Unlike online profiling methods, our offline BO optimizes attention reconstruction error (MSE) on a physics-based proxy task, ensuring rapid convergence. Experiments on HunyuanVideo and Wan2.1-14B demonstrate that DynamicRad pushes the efficiency--quality Pareto frontier, achieving extbf{1.7imes--2.5imes inference speedups} with extbf{over 80% effective sparsity}. In some long-sequence settings, the dynamic mode even matches or exceeds the dense baseline, while mask-aware LoRA further improves long-horizon coherence. Code is available at https://github.com/Adamlong3/DynamicRad.
Related
- Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
- Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement
- Noise-Adaptive Diffusion Sampling for Inverse Problems Without Task-Specific Tuning
- Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
- Remote Sensing Image Super-Resolution for Imbalanced Textures: A Texture-Aware Diffusion Framework
Source: arXiv cs.CV | 2026-04-23