Self-training framework based on multi-granularity local cores for class-imbalanced semi-supervised classification in business applications.

Li, Junnan; Su, Xiaosheng; Wang, Leo; Ma, Jicheng; Xia, Yingjun; Gu, Yuqing; Fu, Shun; Xu, Wenli et al. · Neural Netw · 2026

Where this comes from

Abstract

Self-training methods often suffer from performance degradation in semi-supervised scenarios due to insufficient initial labeled data for training accurate initial classifiers. Recent frameworks like the K-means-based framework for semi-supervised classification (K-means-SSC) and the local cores-based framework for semi-supervised classification (LC-SSC) leverage unlabeled data representatives found by clustering algorithms and predicted by co-labeling or active labeling strategy to improve the initial labeled data to address this. Nevertheless, they still experience the following issues: a) having insufficient improvement for minority class labeled data; b) co-labeling or active labeling strategy for predicting found unlabeled data representatives has a low accuracy for the minority class or a high manual interference degree. These lead to ineffectively improving insufficient labeled data for self-training methods on class-imbalanced semi-supervised data. To overcome these limitations, a self-training framework based on multi-granularity local cores for class-imbalanced semi-supervised classification (MGLC-CISSC) is proposed. Our framework introduces two key contributions: (a) a multi-granularity local core search algorithm (MGLORE) that identifies more fine-grained unlabeled representative data in minority-class regions, and (b) a divide-and-conquer labeling strategy for class-imbalanced semi-supervised data (DCLSCISS) that accurately predicts labels for these data representatives with minimal manual effort. Experiments on class-imbalanced semi-supervised benchmark datasets from industrial applications have demonstrated that MGLC-CISSC significantly outperforms state-of-the-art solutions in enhancing four self-training methods for training two classifiers. The results conclusively show that the proposed MGLC-CISSC more effectively mitigates the labeled-data insufficiency problem, especially for minority classes on class-imbalanced semi-supervised data, advancing the practicality of self-training in class-imbalanced semi-supervised environments of real-world business applications.