Research
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
arXiv:2603.23488v2 Announce Type: replace Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is
arXiv:2603.23488v2 Announce Type: replace Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is not necessary: one view is enough. We present OVIE, trained entirely on unpaired internet images. We leverage a monocular depth estimator as a geometric scaffold at training time: we lift a source image into 3D, apply a sampled camera transformation, and project to obtain a pseudo-target view. To handle disocclusions, we introduce a masked training formulation that restricts geometric, perceptual, and textural losses to valid regions, enabling training on 30 million uncurated images. At inference, OVIE is geometry-free, requiring no depth estimator or 3D representation. Trained exclusively on in-the-wild images, OVIE outperforms prior methods in a zero-shot setting, while being 600x faster than the second-best baseline. Code and models are publicly available at https://github.com/AdrienRR/ovie.
Related
- Novel View Synthesis as Video Completion
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Unfolding 3D Gaussian Splatting via Iterative Gaussian Synopsis
- PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting
- Self-Improving 4D Perception via Self-Distillation
- 3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
Source: arXiv cs.CV | 2026-04-15