C3Net: A cross-modal collaborative calibration of features for object detection using frames and events.

Chen, Yunhua; Zhong, Jinyu; Guo, Yihao; Xie, Zequan; Xiao, Jinsheng; Chen, Pinghua · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Object detection by fusing RGB frames and event streams is challenging due to their inherent heterogeneity and significant statistical disparities, which often lead to suboptimal fusion in existing methods. To address this, we introduce C3Net, a novel framework built upon a paradigm shift from direct feature merging to Collaborative Calibration. First, we propose an Adaptive Balancing Time Surface (ABTS) to generate motion-robust event representations by mitigating spatial inconsistencies caused by varying object velocities. Second, the core Cross-Modal Feature Collaborative Calibration Module (CM-FCCM) performs mutual calibration of RGB and event features across channel and spatial dimensions, reducing modality discrepancies before fusion; the calibrated features are then fed back to the respective backbones for enriched feature learning. Finally, an Adaptive Channel Fusion Module (ACFM) dynamically integrates the modalities based on channel-wise confidence. Extensive experiments on PKU-DAVIS-SOD, DSEC-MOD, and PKU-DDD17-CAR datasets demonstrate that C3Net achieves state-of-the-art performance, showcasing its superior ability to leverage the complementary strengths of frames and events.

Medical subject headings