RegTrack: Simplicity Beneath Complexity in Robust Multi-Modal 3D Multi-Object Tracking.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42546008.
- Also identified by DOI 10.1109/TPAMI.2026.3719678.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Existing 3D multi-object tracking (MOT) methods often trade efficiency and generalizability for robustness, as they typically rely on complex association metrics derived from multi-modal architectures or class-specific motion priors. Challenging the common belief that greater complexity necessarily leads to stronger robustness, we propose a robust, efficient, and generalizable method for multi-modal 3D MOT, dubbed RegTrack. Inspired by Yang-Mills gauge theory, RegTrack formulates multi-modal 3D MOT as motion-compensated representation learning. Under this analogy, point-cloud object representations are viewed as matter fields, while inter-frame object motions are regarded as local variations. Geometric cues are modeled as gauge fields to adaptively compensate for such variations, and a pretrained image representation space serves as a globally invariant physical law to guide the compensation process. In this way, the resulting motion-compensated point-cloud representations, viewed as observables, are encouraged to remain consistent for the same object across frames while preserving discriminability among different objects. Their pairwise similarities thus provide a simple yet robust association metric. Specifically, RegTrack is built upon a unified tri-cue encoder (UTEnc), which consists of a local-global point cloud encoder (LG-PEnc), a mixture-of-experts-based geometry encoder (MoE-GEnc), and a frozen image encoder derived from a pretrained vision-language model. LG-PEnc efficiently encodes the spatial-structural information of object point clouds to generate foundational representations. MoE-GEnc interacts with LG-PEnc to model inter-frame geometric relationships and adaptively compensate for motion-induced representation variations without relying on class-specific priors. The frozen image encoder is used only during training to provide a stable representation space for supervising the compensation process, and is discarded during inference. As a result, RegTrack achieves robust, efficient, and generalizable inference using only point-cloud inputs, with merely 2.67 M parameters. Extensive experiments on KITTI and nuScenes demonstrate that RegTrack outperforms its thirty-five competitors.