Research
A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources
arXiv:2608.13183v1 Announce Type: new Abstract: Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current sca
arXiv:2608.13183v1 Announce Type: new Abstract: Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant benefits arise when effective models can be obtained with fewer resources. To better understand how self-supervised learning (SSL) objectives behave under resource constraints, we conduct a controlled study of image and video SSL objectives under matched data, architecture, and compute budgets. We compare contrastive, reconstruction, feature-prediction, and diffusion objectives and evaluate both standalone and jointly trained image-video SSL formulations across a diverse set of image and video understanding tasks. Our results show that DINOv2-style pretraining consistently provides the strongest overall performance under limited resources. Furthermore, combining DINOv2 with video SSL objectives such as VideoMAE substantially improves image classification and segmentation performance, but degrades video tracking and camera-pose estimation performance, revealing an important tradeoff between semantic and geometric representation learning. These findings suggest that combining image and video SSL objectives can be beneficial in resource-limited settings, while highlighting the need for improved methods that better balance semantic, temporal, and geometric supervision.
Related
- Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
- Boosting Visual Instruction Tuning with Self-Supervised Guidance
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
- Self-supervised pretraining for an iterative image size agnostic vision transformer
Source: arXiv cs.CV | 2026-08-14