Model Releases
RegionDet: A Benchmark for Region Detection Beyond Object Instances
arXiv:2608.06850v1 Announce Type: new Abstract: Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bo
arXiv:2608.06850v1 Announce Type: new Abstract: Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bounded object instances. However, many visual targets in real-world scenarios are not individual objects, but regions defined by visual states, scene context, object relations, and human activities, such as construction areas, damaged road regions, queues, group conversations, and vendor regions. Existing detection benchmarks are mainly built around object instances, providing limited support for systematically evaluating such region targets. To address this gap, we introduce Region Detection, a task that extends conventional object detection beyond object instances, and construct RegionDet, a benchmark for region target localization. RegionDet contains eight region categories, including Construction, Crossing, Damage, Queuing, Talking, Vendor, Waiting, and Walking, with COCO-style bounding-box annotations and evaluation protocols. We systematically evaluate representative closed-set and zero-shot/open-vocabulary detectors on RegionDet. Results show that closed-set detectors can partially learn region-level patterns under supervision, while zero-shot/open-vocabulary detectors struggle severely, revealing the strong object-centric bias of current vision-language detectors. Further analyses highlight key challenges in Region Detection, including weak boundary cues, strong context dependency, and insufficient relation-level region understanding. The RegionDet will be released.
Related
- ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness
- ContextShift: A Controlled Benchmark for Context Dependence in Object Detection
- FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
- LV-OSD: Language-Vision-Complementary Open-Set Object Detection
- Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
Source: arXiv cs.CV | 2026-08-10