Research
HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
arXiv:2604.10772v1 Announce Type: new Abstract: 3D layout generation and editing play a crucial role in Embodied AI and immersive VR interaction. However, manual creation requires tedious labor, while
arXiv:2604.10772v1 Announce Type: new Abstract: 3D layout generation and editing play a crucial role in Embodied AI and immersive VR interaction. However, manual creation requires tedious labor, while data-driven generation often lacks diversity. The emergence of large models introduces new possibilities for 3D scene synthesis. We present HOG-Layout that enables text-driven hierarchical scene generation, optimization and real-time scene editing with large language models (LLMs) and vision-language models (VLMs). HOG-Layout improves scene semantic consistency and plausibility through retrieval-augmented generation (RAG) technology, incorporates an optimization module to enhance physical consistency, and adopts a hierarchical representation to enhance inference and optimization, achieving real-time editing. Experimental results demonstrate that HOG-Layout produces more reasonable environments compared with existing baselines, while supporting fast and intuitive scene editing.
Related
- LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
- WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
- 2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
- Image-Guided Geometric Stylization of 3D Meshes
Source: arXiv cs.CV | 2026-04-14