Research
VDLF-Net: Variational Feature Fusion for Adaptive and Few-Shot Visual Learning
arXiv:2604.23641v1 Announce Type: new Abstract: This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature
arXiv:2604.23641v1 Announce Type: new Abstract: This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature maps, while ell_2-normalized embeddings from the gated maps contribute toward supervised classification or episodic few-shot prediction. Under standard CIFAR-100 and Mini-ImageNet protocols, VDLF-Net demonstrates an improved performance over ResNet-50 Enhanced, VGG-16, Prototypical Networks, and Matching Networks. Extensive ablations show that removing the fine-resolution scale has the greatest impact on VDLF-Net's performance. At the same time, KL and reconstruction at the chosen alpha pose a minor performance reduction, demonstrating that performance gains over classical episodic baselines mainly originate from the full VDLF-Net architecture and training strategy.
Related
- A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction
- Weak-to-Strong Knowledge Distillation Accelerates Visual Learning
- Light 'em Up: Enabling Few-Shot Low-Light 3D Gaussian Splatting with Multi-Scale Explicit Retinex Illumination Decoupling
- Chaos-Enhanced Prototypical Networks for Few-Shot Medical Image Classification
Source: arXiv cs.CV | 2026-04-28