Local Ai

DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models

arXiv:2608.23850v1 Announce Type: new Abstract: Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-vi

DGX agentpaper
local-aiarxiv-cs-cv

arXiv:2608.23850v1 Announce Type: new Abstract: Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.

Related

Source: arXiv cs.CV | 2026-08-26

Loading related sources…