Model Releases
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
arXiv:2607.14595v2 Announce Type: replace Abstract: Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational
arXiv:2607.14595v2 Announce Type: replace Abstract: Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational costs. Existing parameter-efficient fine-tuning (PEFT) methods have two critical flaws on billion-scale models: they still require substantial trainable parameters, and reward-based training suffers from noise-induced optimization instability in condition-guided tasks. We propose MagicPrompt, a lightweight framework that achieves extreme parameter efficiency and stable reward optimization. It first adopts Attention-Embedded Prompt Tuning, which steers generation via lightweight soft prompts with orders of magnitude fewer parameters while preserving pre-trained knowledge. It further introduces Dual-Space Reward Feedback Optimization, which uses self-supervised latent objectives to improve condition-guided reward training. Experiments show MagicPrompt reaches competitive performance with less than 1% trainable parameters and notably reduces training costs.
Related
- Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
Source: arXiv cs.CV | 2026-07-23