Model Releases
ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation
arXiv:2608.21194v1 Announce Type: new Abstract: Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches enco
arXiv:2608.21194v1 Announce Type: new Abstract: Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6% in accuracy while utilizing 590imes fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.
Related
- Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models
- MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
- From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
- Spike-NVPT: Learning Robust Visual Prompts via Bio-Inspired Temporal Filtering and Discretization
Source: arXiv cs.CV | 2026-08-24