Tutorials
Class Geometry as Supervision for Sample-Efficient Open-World Detection
arXiv:2608.12698v1 Announce Type: new Abstract: Open-world object detection requires models to recognize known categories, reject unfamiliar objects, and incorporate new classes over time. This is esp
arXiv:2608.12698v1 Announce Type: new Abstract: Open-world object detection requires models to recognize known categories, reject unfamiliar objects, and incorporate new classes over time. This is especially challenging in scarce-data settings such as biomedical and scientific imaging, where rare categories may have only a few annotated examples and fine-grained classes differ by subtle morphology. Prototype-based detectors are natural for this regime, but they typically learn class prototypes as independent anchors, ignoring relational structure among classes. We propose class-geometry supervision (CGS), a general framework that constrains learned prototype or class-representation spaces to preserve visual or semantic class dissimilarities estimated from training data. CGS introduces a dissimilarity-preserving objective that aligns pairwise distances among learned class representations with a target class-geometry matrix while retaining the standard task loss. We instantiate the same objective across prototype recognition, few-shot biomedical object detection, open-set detection, novel-class insertion, and OWOD adaptation on COCO. Experiments show that CGS improves sample efficiency in recognition and ova detection, substantially strengthens novel-class insertion, and improves unknown recall on COCO while retaining much of the known-class detection performance. Ablations show that meaningful visual geometry provides the most reliable gains, while random geometry can help novel separation but is less consistent for few-shot detection. These results suggest that relational class geometry is an effective supervisory signal for building calibrated and extensible open-world detectors under limited supervision.
Related
- Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
- VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction
- PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought
Source: arXiv cs.CV | 2026-08-14