MANNERS: A strategy for representation learning in multivariate datasets with high proportions of missing data.
Where this comes from
- Record sourced from PubMed, PMID 42453694.
- Also identified by DOI 10.1016/j.patter.2026.101543 and PMC identifier 13366519.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Missing data present a major challenge for deep learning, and various imputation techniques exist. However, imputation quality generally decreases as missing data rates increase. In time-series data from electronic health records, missing rates for individual parameters up to 99% can be observed, posing a problem for mere imputation. In addition, underlying patterns between missing and observed data can contain valuable information that can be learned by deep learning models. In this work, we propose a strategy called missing adjusted normalization and nullity encoding representation strategy (MANNERS), which can be applied to training pipelines for representation learning. MANNERS encodes missingness, masks loss evaluation at missing data points, and applies rebalancing such that the signal from variables with high missing rates is not lost. We evaluated MANNERS on reconstruction and downstream classification, regression, and synthetic data generation tasks. We showed performance improvements in the presence of very high missing rates compared with state-of-the-art imputation-only techniques.