Language-4D Cross-Boosting for Generalized Zero-Shot 6DoF Tracking and 3D Reconstruction.

Sun, Jingtao; Wang, Yaonan; Zhao, Jiawen; Zhang, Yike; Liu, Min; Shou, Mike Zheng · IEEE Trans Pattern Anal Mach Intell · 2026

basic_science · Level V

Where this comes from

Abstract

A more universal category-level pose tracking and reconstruction approach should be robust to both seen and unseen classes. However, the performance of existing category-level algorithms degrades for classes that are unseen in the training set. To learn a discriminative model that can achieve strong performance in Generalized Zero-Shot Learning (GZSL) settings, we introduce L4D-Track++, a unified framework for zero-shot category-level pose tracking and shape reconstruction that leverages multi-level language captions. To address the bias problem, namely the out-of-distribution gap between source (seen) and target (unseen) domains, we first propose a Multi-Level Language Conditional Feature Generation module to produce synthetic pairwise embeddings for unseen classes. Furthermore, to increase inter-class distance, decrease intra-class distance, and bridge the gap between seen and unseen category domains, we propose a language-guided multi-modal fusion strategy to construct the final fused inter-frame pairwise embeddings for downstream tracking and reconstruction tasks by taking both real and synthetic pairwise embeddings as inputs. This strategy consists of two main components: Category-Related Mutual Information-based Optimization and Modality-Related Conditional Mutual Information Maximization. Experimental results on five public datasets demonstrate that our method is effective and achieves significant improvements in GZSL performance.