ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models
arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pa