Agents
Direct and Adaptable Mesh-Gaussian Scene Reconstruction from Multi-View Images
arXiv:2405.06945v4 Announce Type: replace Abstract: Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential fo
arXiv:2405.06945v4 Announce Type: replace Abstract: Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential for maintaining high-fidelity real-to-sim environments for embodied intelligence, where local changes should be incorporated without complete reconstruction. Existing neural surface reconstruction and 3DGS-to-mesh pipelines often learn geometry indirectly or separate geometry construction from appearance modeling. This separation introduces optimization redundancy and makes local geometry or appearance updates expensive. We propose an end-to-end mesh-Gaussian scene representation that binds 3D Gaussians to mesh faces and uses differentiable 3DGS rendering for photometric supervision. This design provides a direct information pathway for jointly learning explicit geometry and renderable appearance. Experiments on indoor and outdoor scenes demonstrate improved efficiency and rendering quality while preserving high-quality surface reconstruction. The explicit mesh also enables mesh-based manipulation, and the coupled representation adapts efficiently to local scene modifications. These properties support scalable visual scene modeling and the efficient maintenance of real-to-sim environments for embodied-agent training and evaluation.
Source: arXiv cs.CV | 2026-08-04