Tutorials

Generalizable Operating Room Expert with Multimodal Enhancement

arXiv:2508.08199v2 Announce Type: replace Abstract: Precise spatial modeling in the operating room (OR) is essential for intraoperative awareness, hazard avoidance, and surgical decision-making. Altho

DGX agentpaper
tutorialsarxiv-cs-cv

arXiv:2508.08199v2 Announce Type: replace Abstract: Precise spatial modeling in the operating room (OR) is essential for intraoperative awareness, hazard avoidance, and surgical decision-making. Although existing approaches exploit multimodal data to learn spatial relationships, many depend on sensing modalities that are difficult to deploy in real clinical environments and remain limited in explicit 3D reasoning under constrained sensing conditions. Meanwhile, models trained primarily on readily available 2D data often fail to capture the fine-grained geometric and semantic structure of complex OR scenes. To address these limitations, we introduce extbf{OR-Expert}, a large vision-language framework for 3D spatial reasoning with RGB-only inference. OR-Expert internally derives depth, panoptic segmentation, and point-cloud cues from RGB images and encodes them as structured spatial representations. Its Spatial-Enhanced Feature Fusion Block aligns these pseudo-modalities with RGB and textual features in a shared token space, enabling joint semantic, geometric, and language reasoning. The unified end-to-end MLLM therefore supports detailed spatial understanding without requiring external depth, segmentation, or point-cloud sensors, or additional expert annotations at inference time. Experiments on multiple operating-room benchmarks demonstrate that OR-Expert achieves state-of-the-art performance and generalizes effectively to unseen surgical scenes and downstream spatial reasoning tasks.

Related

Source: arXiv cs.CV | 2026-08-14

Loading related sources…