Training Dataset Curation by L<sub>1</sub>-Norm Principal-Component Analysis for Support Vector Machines.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40402706.
- Also identified by DOI 10.1109/TNNLS.2025.3568694.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Support vector machines (SVMs) have been the learning model of choice in numerous classification applications. While SVMs are widely successful in real-world deployments, they remain susceptible to mislabeled examples in training datasets where the presence of few faults can severely affect decision boundaries, thereby affecting the model's performance on unseen data. In this brief, we develop and describe in implementation detail a novel method based on $L_{1}$ -norm principal-component data analysis and geometry that aims to filter out atypical data instances on a class-by-class basis before the training phase of SVMs and thus provide the classifier with robust support-vector candidates for making classification boundaries. The proposed dataset curation method is entirely data-driven (touch-free), unsupervised, and computationally efficient. Extensive experimental studies on real datasets included in this brief illustrate the $L_{1}$ -norm curation method and demonstrate its efficacy in protecting SVM models from data faults during learning.