Research
Improving Diversity in Black-box Few-shot Knowledge Distillation
arXiv:2604.25795v1 Announce Type: new Abstract: Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacri
arXiv:2604.25795v1 Announce Type: new Abstract: Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacrifice in performance. However, most KD methods require a large training set and internal access to the teacher, which are rarely available due to various restrictions. These challenges have originated a more practical setting known as black-box few-shot KD, where the student is trained with few images and a black-box teacher. Recent approaches typically generate additional synthetic images but lack an active strategy to promote their diversity, a crucial factor for student learning. To address these problems, we propose a novel training scheme for generative adversarial networks, where we adaptively select high-confidence images under the teacher's supervision and introduce them to the adversarial learning on-the-fly. Our approach helps expand and improve the diversity of the distillation set, significantly boosting student accuracy. Through extensive experiments, we achieve state-of-the-art results among other few-shot KD methods on seven image datasets. The code is available at https://github.com/votrinhan88/divbfkd.
Related
- Diverse Image Priors for Black-box Data-free Knowledge Distillation
- Weak-to-Strong Knowledge Distillation Accelerates Visual Learning
- Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
- Weighted Knowledge Distillation for Semi-Supervised Segmentation of Maxillary Sinus in Panoramic X-ray Images
- Protecting Language Models Against Unauthorized Distillation through Trace Rewriting
Source: arXiv cs.CV | 2026-04-29