Unified Video Dense Prediction from Disjoint Data
arXiv:2607.21592v1 Announce Type: new Abstract: Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragment