MVDG: Efficient Multi-view 3D Disambiguation on Unconstrained Real-World Images
Hanyuan Xiao, Gonglin Chen, Haolin Xiong, Wenbin Teng, Haiwei Chen, Yajie Zhao
Abstract
Illusory matches between distinct yet visually similar 3D surfaces--doppelgangers--remain a fundamental obstacle for large-scale, in-the-wild 3D reconstruction and visual localization. Prior work mitigates this issue with pairwise classifiers, but this design limits multi-view contextual reasoning and incurs O(n2) inference complexity for downstream structure-from-motion (SfM). We present MVDG, a scalable multi-view disambiguation framework built on the 3D foundation model VGGT, which jointly reasons over an arbitrary number of multiview images. By incorporating 3D-aware multi-view features, our method reduces dependence on pairwise comparisons by encoding and decoding views in a single pass. We further observe that direct multi-view fine-tuning of VGGT can be unstable under noisy supervision; motivated by label ambiguity in Doppelgangers, we construct a pseudo-pairwise training set from AerialMegaDepth and show that fine-tuning on sampled subsets yields stable optimization and strong generalization to held-out scenes. Finally, because full SfM evaluation (even with faster pipelines such as GLOMAP) remains expensive, we process a pseudo-pairwise dataset for efficient validation; we derive a predictive relationship between regular SfM metrics and the classification accuracy on this pseudo-pairwise test. Experiments show that our method achieves comparable pairwise accuracy while improving both SfM accuracy and inference speed over baselines.
Create a lesson
Related papers
Moore, Escher, Penrose: A Conformal Golden Braid
Sophia Feldman, Assaf Shocher
Sphere Encoder 2
Kaiyu Yue, Sean McLeish, Ruchit Rawal et al.
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
ROWBench: Do Video Models Render What the Program Specifies?
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai et al.
Embedding Prediction Helps Image Generation
Sihan Xu, Ji Xie, Zilin Wang et al.
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.