Model Releases
FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
arXiv:2608.12683v1 Announce Type: cross Abstract: Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire o
arXiv:2608.12683v1 Announce Type: cross Abstract: Embodied agents must often identify and interact with objects based on their function rather than their identity, requiring them to actively acquire observations that reveal discriminative functional evidence. Existing affordance grounding methods operate from fixed viewpoints and lack mechanisms for deciding where to look when functional cues are occluded or incomplete. We introduce Active Functional Affordance Grounding, a new task in which an agent sequentially explores a scene to identify and spatially ground an object satisfying a functional query. To address this problem, we propose FUSE, an adaptive semantic-geometric evidence acquisition framework that combines explicit uncertainty-driven exploration with a learned amortized planner to efficiently select informative viewpoints. We further introduce a Habitat-based benchmark for evaluating active functional grounding. Experiments show that FUSE achieves the highest observed non-oracle grounding performance while reducing computation by 1.33x relative to fully explicit exploration, and remains effective across multiple affordance knowledge sources.
Related
- GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
- CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
- EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
- Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
Source: arXiv cs.CV | 2026-08-14