BLEND-M: cellular deconvolution with personalized DNA methylation references
Abstract
Cellular deconvolution estimates cell type fractions in tissue samples, providing an in-silico alternative to costly and labor-intensive cell sorting techniques. Accurate fraction estimates are essential for DNA methylation (DNAm) analyses, as they enable adjustment for confounding, support risk prediction models, and facilitate cell type-specific (CTS) differential methylation analyses. A key determinant of deconvolution accuracy is the choice of CTS reference profiles, which represent expected DNAm levels in each cell type. Although it is widely recognized that CTS DNAm levels vary substantially across individuals, existing methods assume a shared reference across the population, leading to biased fraction estimates. We introduce BLEND-M, a novel method for deconvolving bulk DNAm data that learns personalized CTS reference profiles for each bulk sample. In addition to accounting for inter-individual variability in CTS DNAm, BLEND-M explicitly models heteroscedasticity across marker CpGs—the fact that DNAm measurements at some CpGs are noisier than others—thereby improving robustness to noisy markers. We establish theoretical guarantees for the accuracy of BLEND-M and demonstrate its improved performance relative to existing methods through extensive benchmarking on realistic simulated and real datasets. We further illustrate its practical utility by applying BLEND-M to develop a risk prediction model for childhood atopic asthma. BLEND-M improves the accuracy and robustness of DNAm deconvolution by accounting for inter-individual variability in CTS DNAm. These advances enhance downstream analyses, including risk prediction, confounding adjustment, and CTS differential methylation studies.
// Source
Authors: Penghui Huang, David G. Peters, Chris McKennan
Institutions: University of Pittsburgh, Department of Health, Magee-Womens Research Institute