HTNet: A self-supervised heterogeneous triple network for multi-modal data.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42413356.
- Also identified by DOI 10.1016/j.neunet.2026.109313.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Current self-supervised learning methods, predominantly built on Siamese architectures, have limited ability to handle small-scale, image-text multi-modal data. This paper introduces HTNet, a self-supervised triple network designed for such multi-modal tasks. Unlike conventional Siamese networks that consist of two identical sub-networks, HTNet incorporates a third, heterogeneous branch, creating a triple network architecture that distills knowledge into a student from both an isomorphic and a heterogeneous teacher. We further propose an adaptive diversity loss function that balances the influence of these structurally different teachers, ensuring the student learns a unified and comprehensive feature representation. Experimental results demonstrate that HTNet achieves strong feature extraction on small-scale datasets including STL-10, CIFAR-10, and CIFAR-100, exhibits excellent generalization in transfer learning after pre-training on Tiny-ImageNet, and outperforms state-of-the-art self-supervised methods on image-text retrieval benchmarks including MS-COCO and Flickr30K.