Temporal and spatial context aware voxel transformer for semantic scene completion.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41780283.
- Also identified by DOI 10.1016/j.neunet.2026.108754.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Semantic Scene Completion (SSC) aims to recover comprehensive 3D structure and semantic understanding from partial visual observations, playing a crucial role in autonomous driving perception. However, current camera-based SSC methods often suffer from insufficient temporal reasoning and unreliable depth estimation, which results in incomplete geometry recovery and ambiguous semantic interpretation. To address these challenges, we develop a method that recovers complete 3D semantics by progressively integrating multi-frame context and depth cues. Starting from input images, contextual and depth-related features are separately extracted. Contextual features are then temporally aligned through temporal-spatial aware mechanisms that maintain coherence under viewpoint variations. Meanwhile, monocular depth priors are probabilistically fused with stereo-derived constraints to refine depth estimates. The refined depth further guides volumetric reconstruction, enabling accurate and stable semantic scene completion. The proposed method is evaluated on SemanticKITTI and SSCBench-KITTI-360, achieving results on par with or surpassing current state-of-the-art methods. The source code is available at: https://www.github.com/djzgroup/SSC.
Medical subject headings
- Semantics
- Imaging, Three-Dimensional