Robust Self-Supervised Monocular Depth Estimation for Endoscopic Soft Tissue Deformation Scenes With Biomechanical Constraints.

Wang, Enpeng; Xu, Jiangchang; Liu, Yueang; Tu, Puxun; Wang, Junfeng; Jiang, Xiaoyi; Chen, Xiaojun · IEEE Trans Image Process · 2026

basic_science · Level V

Where this comes from

Abstract

Self-supervised learning technology has been applied to calculate depth and ego-motion from monocular videos, achieving remarkable performance in various real-world scenarios. Unfortunately, challenges such as specular reflections and soft tissue deformations in endoscopic scenes greatly undermine the performance of these methods, inevitably compromising the accuracy of depth and ego-motion estimation. To address these two problems, we introduce a novel strategy based on image distance transform for robust self-supervised learning for monocular depth estimation, effectively handling specular reflections in endoscopic scenes. Furthermore, we propose a soft tissue deformation constraint based on biomechanical principles, which mitigates the adverse effects of deformed region pixels, ultimately enhancing the model's depth estimation precision. Additionally, our method employs a lightweight architecture ensuring a reduced number of model parameters and faster inference time. Extensive experiments are conducted on both public datasets (SCARED, SERV-CT) and our own datasets to validate the effectiveness of our method. Compared with other SOTA methods, our approach demonstrates comparable accuracy and robustness while ensuring faster inference time. On the SCARED dataset, our approach attains an RMSE of 4.96 mm with only 2.25M model parameters for depth estimation. Especially, experiment results on SERV-CT dataset and our own datasets further demonstrate the model's generalization ability and potential clinical value in computer-assisted surgical navigation.

Medical subject headings