Clustering of temporal gene expression data with mixtures of mixed effects models with a penalized likelihood.
Where this comes from
- Record sourced from PubMed, PMID 30101356.
- Also identified by DOI 10.1093/bioinformatics/bty696 and PMC identifier 6394398.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
MOTIVATION: Clustering algorithms like K-Means and standard Gaussian mixture models (GMM) fail to account for the structure of variability of replicated data or repeated measures over time. Additionally, a priori cluster number assumptions add an additional complexity to the process. Current methods to optimize cluster labels and number can be inaccurate or computationally intensive for temporal gene expression data with this additional variability. RESULTS: An extension to a model-based clustering algorithm is proposed using mixtures of mixed effects polynomial regression models and the EM algorithm with an entropy penalized log-likelihood function (EPEM). The EPEM is used to cluster temporal gene expression data with this additional variability. The addition of random effects in our model decreased the misclassification error when compared to mixtures of fixed effects models or other methods such as K-Means and GMM. Applying our method to microarray data from a fracture healing study revealed distinct temporal patterns of gene expression. AVAILABILITY AND IMPLEMENTATION: https://github.com/darlenelu72/EPEM-GMM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms