BEVFix: Deep feature enhancement for robust 3D object detection.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40505163.
- Also identified by DOI 10.1016/j.neunet.2025.107675.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recent advancements in Bird's Eye View (BEV)-based 3D object detection have highlighted its potential to enhance scene understanding in autonomous driving applications. However, existing BEV-based methods utilizing point clouds for 3D object detection face significant challenges due to inherent sparsity and noise, which often compromise the accuracy of BEV representations. Furthermore, in multimodal 3D object detection, the lack of depth information in images can lead to distortions in the image BEV features generated through view transformations, further leading to inaccuracies in the fused BEV representation. To overcome these limitations, we introduce BEVFix, an innovative end-to-end 3D object detection method designed to refine BEV representations. BEVFix starts by generating a mask based on the point cloud distribution to identify specific regions requiring repair. This is followed by our WaveRefiner, which employs Discrete Wavelet Transform (DWT) for multi-frequency decomposition and utilizes a Feed-Forward Network (FFN) to isolate noise while selectively retaining critical features. These components work synergistically to reduce noise and enhance BEV representations. Experiments on the nuScenes and Waymo datasets demonstrate that BEVFix significantly improves performance, achieving state-of-the-art results. The source code will be publicly available at https://github.com/WenxuanLi-whu/Co-Fix3d.
Medical subject headings
- Imaging, Three-Dimensional
- Neural Networks, Computer