Model Releases
ViSculpt: Visual-Centric Agentic Geometry Editing
arXiv:2608.24169v1 Announce Type: new Abstract: 3D geometry editing is a critical yet labor-intensive part of the graphics pipeline, requiring artists to translate creative intent into precise operati
arXiv:2608.24169v1 Announce Type: new Abstract: 3D geometry editing is a critical yet labor-intensive part of the graphics pipeline, requiring artists to translate creative intent into precise operations in complex professional software. Large language models (LLMs) have shown promise for script-based 3D creation, but script generation is less suited to perception-driven editing of arbitrary existing meshes, where execution must remain visually grounded and untouched regions should be preserved. We present a visual-centric, training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists. Rather than generating scripts or regenerating geometry, our system operates through the Blender GUI: multimodal LLM agents observe the viewport, reason about the current mesh state, and execute localized edits through simulated user interactions. Experiments on a curated benchmark provide initial evidence that this agentic approach can follow natural language instructions, perform representative localized mesh edits, and preserve the overall identity of the input asset. Our results highlight a complementary regime for language-driven 3D editing: direct in-place modification of existing meshes within the native 3D editing workflow. We view this work as an exploratory step toward visual-centric agentic geometry editing in professional graphics software.
Related
- EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation
- StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
- Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Source: arXiv cs.CV | 2026-08-26