A two-stage active cleaning strategy for long-tail label noise.
Where this comes from
- Record sourced from PubMed, PMID 41544499.
- Also identified by DOI 10.1016/j.neunet.2026.108585.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Long-tailed data is ubiquitous in real-world applications, posing significant challenges due to imbalanced class distribution and high levels of label noise. Previous methods to address long-tailed data with label noise often incur high computational and manual costs. To address these challenges, we propose a novel two-stage active label cleaning strategy, which extends active learning beyond traditional label acquisition to efficiently identify and correct mislabeled samples while minimizing annotation cost. Specifically, in the first stage, we propose a Balanced Class-Centered Contrastive Learning (BCCL) to enhance feature representation quality and identify potential label noise within long-tailed datasets. BCCL achieves this through a novel loss function that integrates contrastive learning with weighted average of class centers. The second stage employs an uncertainty-based active learning sorted sampling to retrain on potential label noise samples, focusing on high-uncertainty instances to determine the final noise samples needing to be relabeled. Our two-stage active label cleaning strategy minimizes the amount of data requiring re-annotation, ultimately improving classification performance through iterative re-labeling, while optimizing the use of annotation resources and reducing the annotation workload. Experimental results demonstrate the robustness of our proposed method across varying noise ratios and levels of imbalance, effectively enhancing discriminative capability on noisy data in multiple datasets and achieving superior classification performance on long-tailed data, particularly in high-noise scenarios. In experiments on the CIFAR10-LT dataset under imbalance ratio 10 and symmetric noise 0.6, we significantly outperform the state-of-the-art PCSE with a relative improvement of 5.17%. In addition, on the real-noise long-tail dataset Red Mini-ImageNet under imbalance ratio 100 and noise ratio 0.4, we achieve an accuracy of 38.37%, surpassing existing baselines.
Medical subject headings
- Machine Learning
- Neural Networks, Computer