Model Releases

DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis

arXiv:2603.23924v2 Announce Type: replace Abstract: Text-to-image models often struggle to synthesize correct occlusion relationships among multiple objects, especially in densely overlapping regions.

DGX agentpaper
model-releasesarxiv-cs-cv

arXiv:2603.23924v2 Announce Type: replace Abstract: Text-to-image models often struggle to synthesize correct occlusion relationships among multiple objects, especially in densely overlapping regions. Many training-free layout-guided methods enforce 2D spatial constraints but do not explicitly resolve depth-dependent attention competition, which can cause concept mixing and implausible occlusion. To address this problem, we propose DepthArb, a training-free framework that formulates occlusion generation as attention arbitration within a unified denoising trajectory. DepthArb employs two core occlusion-control mechanisms: Attention Arbitration Modulation suppresses background-object attention within foreground support according to relative depth, while Spatial Compactness Control limits attention dispersion to preserve object coherence. Because interference varies during generation, Occlusion Conflict Estimation constructs a shared spatial conflict field to adaptively weight both objectives. Through a unified spatial-text attention interface, DepthArb operates on U-Net cross-attention and the image-to-text component of MMDiT joint attention without model retraining. We further introduce OcclBench, a benchmark with continuous relative-depth specifications and occlusion-specific evaluation metrics. Experiments on OcclBench and public benchmarks show that DepthArb improves several layout and occlusion metrics over the evaluated baselines while maintaining competitive text-image alignment.

Source: arXiv cs.CV | 2026-08-18

Loading related sources…