How Relation Enrichment Improves Clustering Ensemble Performance: A Second Order Induced Relation View.
Where this comes from
- Record sourced from PubMed, PMID 42447018.
- Also identified by DOI 10.1109/TPAMI.2026.3713520.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The clustering ensemble technique that combines multiple clustering results is an effective strategy to improve the accuracy and robustness of the final clustering. Most of the clustering ensemble methods are developed based on the co-association matrix (CA) that is formed by the frequency that two samples belonging to the same cluster. The CA is directly constructed based on the observed cluster set and is unable to reveal hidden relationships. Then, many methods have been proposed to enrich the CA matrix. Although the improvement of the enrichment methods on the ensemble performance has been witnessed, the mechanism of the improvement is not well studied. In this paper, we explore how relation enrichment strategy improves clustering ensemble performance from the view of a second order induced relation. Firstly, we design a second order induced co-association relation (SoCo), which realizes the idea that two samples may be in the same potential unobserved cluster if they have common neighbors that simultaneously belong to multiple clusters. We then analyze SoCo from three aspects. We study the computational equations of SoCo and CA to reveal the relations they respectively consider to answer what are their differences. We define cluster ε-Conflict and cluster ε-Harmony to answer whether SoCo is more beneficial for clustering than CA. We give the expectation and variance of SoCo and CA to answer how SoCo improves CA. Finally, based on SoCo, a clustering ensemble method (CE-SoCo) is developed. Experimental analyses on sixteen data sets including five types of data show that the CE-SoCo obtains excellent clustering ensemble performance compared with the other seventeen representative clustering ensemble methods.