Safety
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
arXiv:2604.15770v1 Announce Type: new Abstract: Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel leve
arXiv:2604.15770v1 Announce Type: new Abstract: Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existing representations struggle to jointly satisfy these requirements, and densely propagating pixel-wise semantics to 3D often results in substantial redundancy, leading to inefficient storage and querying in large-scale scenes. To address these challenges, we present PLAF, a Pixel-wise Language-Aligned Feature extraction framework that enables dense and accurate semantic alignment in 2D without sacrificing open-vocabulary expressiveness. Building upon this representation, we further design an efficient semantic storage and querying scheme that significantly reduces redundancy across both 2D and 3D domains. Experimental results show that PLAF provides a strong semantic foundation for accurate and efficient open-vocabulary 3D scene understanding. The codes are publicly available at https://github.com/RockWenJJ/PLAF.
Related
- OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
- SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
- UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation
- Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
- SAM3-I: Segment Anything with Instructions
Source: arXiv cs.CV | 2026-04-20