T-Rex2++: Towards Generic Object Perception Via Text-Visual Prompt Synergy.

Jiang, Qing; Li, Feng; Zeng, Zhaoyang; Ren, Tianhe; Liu, Shilong; Zhang, Lei · IEEE Trans Pattern Anal Mach Intell · 2026

Where this comes from

Abstract

We present T-Rex2++, a unified and highly practical framework for generic open-set object perception, encompassing both object detection and instance segmentation. Previous methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex object representation due to data scarcity and descriptive limitations. Conversely, visual prompts excel in depicting novel objects through concrete visual examples, but fall short in conveying the abstract concept of objects as effectively as text prompts. Recognizing these complementary strengths, we introduce a text-visual synergy mechanism that aligns both modalities within a single feature space via contrastive learning. Crucially, T-Rex2++ advances beyond the passive perception paradigm of its predecessor by introducing a novel Universal Prompt. This learnable component models generic objectness, empowering the system to autonomously discover and localize arbitrary objects without any user-provided cues, thereby closing the loop between human-guided interaction and fully automatic perception. Furthermore, we extend the synergy verification to the pixel level by integrating a zero-shot instance segmentation module, demonstrating that our contrastive alignment generalizes robustly to fine-grained masks. Comprehensive experiments demonstrate that T-Rex2++ exhibits strong zero-shot object perception capabilities across a wide spectrum of scenarios, validating T-Rex2++ as a versatile foundation for generic object perception.