Research
DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
arXiv:2608.20759v1 Announce Type: new Abstract: Single-image 3D human reconstruction often suffers from over-smoothed textures and geometric inconsistencies. While diffusion models improve generative
arXiv:2608.20759v1 Announce Type: new Abstract: Single-image 3D human reconstruction often suffers from over-smoothed textures and geometric inconsistencies. While diffusion models improve generative quality, their reliance on multi-view synthesis prior to 3D reconstruction is computationally expensive and prone to view inconsistency. We propose DiGS-Avatar, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design. To capture accurate spatial structure, we introduce a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student. Treating this inferred latent as a robust structural skeleton, our method injects high-level semantic features to accurately recover fine textural details without disrupting spatial integrity. The refined representation is then decoded into 3D Gaussian primitives. Extensive experiments demonstrate that DiGS-Avatar achieves state-of-the-art or highly competitive visual fidelity and zero-shot generalization, while reconstructing a fully animatable 3D avatar in just 0.71 seconds. Code is available at https://github.com/KLMAV-CUC/DiGS-Avatar.
Related
- Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation
- Image-Guided Geometric Stylization of 3D Meshes
- Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models
- Reconstruction of a 3D wireframe from a single line drawing via generative depth estimation
Source: arXiv cs.CV | 2026-08-24