Graph-driven contextual synergy network for robust 3D object detection.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42561876.
- Also identified by DOI 10.1016/j.neunet.2026.109451.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
3D object detection stands as a pivotal task in scene understanding. However, two primary bottlenecks constrain current methodologies: semantic ambiguity arising from spatial misalignment during cross-modal fusion, and the inadequate contextual representation of individual candidate points within complex scenes. To address these challenges, this paper presents a Graph-driven Contextual Synergy Network (GCS3D), which is designed to systematically enhance point representations across both semantic and geometric dimensions. Specifically, the proposed method incorporates a Semantic Representation Rectification (G-SRR) module for cross-modal representation enhancement. By performing region-level semantic aggregation based on 3D neighborhoods to mitigate projection bias, this module achieves robust cross-modal fusion through a Spatial-aware Gating Mechanism (SGM) that adaptively regulates visual feature injection. Regarding contextual correlation modeling, the framework introduces a Graph-guided Geometric Consistency Interaction (G-GCI) module. By constructing a local topology graph among anchors and executing position-aware feature interaction, this module facilitates the aggregation of complementary neighborhood information, thereby bolstering the feature consistency and discriminability of anchor representations. Furthermore, a Spatial-Scale Aware Assigner (SSA-Assigner) is utilized to dynamically allocate supervision signals based on prediction quality, fully exploiting the performance potential inherent in the enhanced anchor representations. Extensive experiments on the SUN RGB-D and ScanNet V2 datasets demonstrate that GCS3D achieves superior results with mAP@0.25 scores of 70.39 and 73.86 respectively, validating the effectiveness and robustness of the proposed strategy in complex indoor scenes.