Research
ReLIC-SGG: Relation Lattice Completion for Open-Vocabulary Scene Graph Generation
arXiv:2604.22546v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing method
arXiv:2604.22546v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing methods usually treat annotated triplets as positives and all unannotated object-pair relations as negatives. However, scene graph annotations are inherently incomplete: many valid relations are missing, and the same interaction can be described at different granularities, e.g., extit{on}, extit{standing on}, extit{resting on}, and extit{supported by}. This issue becomes more severe in open-vocabulary SGG due to the much larger relation space. We propose extbf{ReLIC-SGG}, a relation-incompleteness-aware framework that treats unannotated relations as latent variables rather than definite negatives. ReLIC-SGG builds a semantic relation lattice to model similarity, entailment, and contradiction among open-vocabulary predicates, and uses it to infer missing positive relations from visual-language compatibility, graph context, and semantic consistency. A positive-unlabeled graph learning objective further reduces false-negative supervision, while lattice-guided decoding produces compact and semantically consistent scene graphs. Experiments on conventional, open-vocabulary, and panoptic SGG benchmarks show that ReLIC-SGG improves rare and unseen predicate recognition and better recovers missing relations.
Related
- CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation
- Frequency-guided Multi-level Reasoning for Scene Graph Generation in Video
- ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining
- Cross-Attentive Multiview Fusion of Vision-Language Embeddings
Source: arXiv cs.CV | 2026-04-27