Multi-Source Temporal-Depth fusion for robust end-to-End visual odometry.

Zhang, Sihang; Cao, Congqi; Gao, Qiang; Liu, Ganchao · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

End-to-end visual odometry models have recently achieved localization accuracy on par with conventional techniques, while effectively reducing the occurrence of catastrophic failures. However, the relevant models cannot leverage the complete time-series data for pose adjustment and optimization. Moreover, these models are limited to using joint depth prediction tasks merely as a means of scale constraint, lacking effective utilization of depth information. In this paper, we propose an end-to-end multi-source visual odometry (MVO) model that dynamically integrates the key components of hybrid visual odometry pipelines into a unified, learnable deep framework. Specifically, we propose TimePoseNet to model the mapping relationship from time to pose, capturing temporal dependencies across the entire sequence. Additionally, a wavelet convolutional attention mechanism is employed to extract global depth information from the depth map, which is then directly embedded into the pose features to dynamically constrain scale ambiguity. Furthermore, temporal and depth cues are jointly incorporated into the post-processing stage of pose estimation. The proposed method attains state-of-the-art performance on both the KITTI benchmark and the newly introduced UAV-2025 dataset, while preserving computational efficiency during inference.

Medical subject headings