Uncertainty-constrained fusion of single-view and multi-view depth estimation for AR virtual-real occlusion.

Liu, Jia; Ding, Shuai; Li, Yongze; Wang, Bin; Wei, Lina; Chen, Dapeng · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Virtual-real occlusion is a critical challenge in augmented reality (AR). In complex scenes, the accuracy and stability of occlusion handling between virtual objects and the physical world strongly affect interaction realism and user immersion. Depth-based methods are attractive for real-time AR due to their computational efficiency and adaptability. However, this approach heavily relies on highly accurate depth estimation. Methods based on hardware sensors are often difficult to deploy widely, while depth estimation algorithms are prone to performance degradation in complex interactive scenarios. Therefore, we design a depth estimation algorithm specifically suited for such environments. We design a hybrid depth estimation algorithm that integrates single-view and multi-view cues. The proposed model augments multi-view cost volumes with features from a single-view encoder and introduces a Bayesian convolution-based uncertainty module to suppress low-confidence predictions, particularly near boundaries and occlusions. A dynamic mask generated by the single-view branch occludes moving regions in non-target frames, reducing the negative impact of dynamics during cost-volume construction. On ScanNet v2, our method improves over the multi-view baseline SimpleRecon with an overall reduction in AbsRel by about 23.9% and RMSE by 28%, while δ<sub>1</sub> increases by 5.43 percentage points. In addition, the boundary intersection over union (IoU) in AR occlusion assessment has been improved by 7.13%, indicating more reliable depth and sharper occlusion handling.

Medical subject headings