Jana T. Winterstein, Lukas Heinlein, Günter Raddatz +34q-bio.GN cs.LG
DNA methylation provides a stable record of cellular identity, capturing epigenetic programs that distinguish specialized cell states despite a shared genome. Because malignant transformation and tumour progression are accompanied by extensive epigenetic remodeling, we hypothesized that the methylome of melanocytic lesions contains biologically and clinically relevant information for both diagnosis and disease progression. In a cohort of 1,001 tissue samples prospectively collected across eight German university hospitals profiled using Illumina Infinium MethylationEPIC arrays, we compared machine-learning models based on selected Cytosine phosphate Guanine (CpG) methylation sites with models incorporating biology-guided features, including epigenetic age acceleration, cell type composition and copy-number variation burden. In an external test set, the best diagnostic classifier was CpG-based and distinguished melanocytic nevi, noninvasive melanoma and invasive melanoma with a macro-averaged area under the receiver operating characteristic curve of 0.919 (95% CI: 0.878 to 0.952). Notably, across CpGs most strongly hyper- and hypomethylated between NV and IM, NIM showed an intermediate methylation profile, providing a molecular correlate of its diagnostic complexity. The best model for clinically relevant treatment group prediction, with AJCC stages grouped according to guideline-based management recommendations, relied on biology-guided features and achieved a macro-averaged mean absolute error of 0.627 (95% CI: 0.477 to 0.808). Together, these findings demonstrate that methylation-based models can capture both diagnostic identity and clinically relevant disease stratification, supporting DNA methylation as a promising biomarker for further validation and potential clinical translation.
Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology. Methods that rely on epigenetic marks such as DNA methylation typically operate on aggregated methylation estimates, discarding the pattern-level information carried by individual DNA reads. Existing read-level approaches that exploit this information are scarce, and all remain restricted to few-class settings; scaling them further is an open problem because, at scale, non-discriminative reads dominate and hard labels conflict with the many-to-many mapping between methylation patterns and cell types, preventing classifier convergence. To overcome this, we propose data-driven soft labels that estimate the conditional cell-type distribution for each read, and integrate this scheme into Syto, a new modular framework for read-level classification-based deconvolution. On a whole-body atlas of 39 human cell types, Syto reduces MSE by 2.56$\times$ over SoTA, with gains transferring to an out-of-distribution dataset spanning 16 tissues. Syto lays the foundation for modeling increasingly large cell-type panels, with improved applications in biology and healthcare. The proposed soft-labeling scheme is further translatable to any setting with a many-to-many signal-to-label mapping.
Chandan Gupta, Syed Haider, Pietro Liòq-bio.QM cs.LG
DNA methylation (DNAm) serves as one of the most robust molecular biomarkers of biological aging. While conventional epigenetic clocks accurately predict chronological age from high-dimensional CpG profiles, they treat aging as a static regression task, meaning they can only output a single score rather than simulating how an entire profile continuously changes over time. To reconstruct these continuous dynamics, we frame lifelong human epigenetic aging as a trajectory inference problem across discrete age snapshots derived from widely available cross-sectional data. We introduce a two-stage computational pipeline: first, an age-regularized Variational Autoencoder (VAE) maps high-dimensional CpG profiles onto a chronologically ordered latent manifold while preserving a generative decoder bridge back to the original methylation space. Second, we model the continuous movement across this latent space via Regularized Unbalanced Optimal Transport (RUOT) that unifies deterministic drift, random diffusion, and non-conservative mass changes. By resolving this RUOT formulation using the DeepRUOT framework, our model fluidly accommodates population-level density shifts like survivorship bias and cellular attrition without requiring rigid biological priors. Evaluated on a large-scale, 80-year pan-tissue dataset, our model demonstrates robust distribution interpolation and uncovers a prominent late-life surge in the learned growth field that mathematically captures the variance expansion driven by stochastic epigenetic drift. Finally, by decoding continuous latent paths back to individual CpG sites, we reconstruct and empirically verify distinct biological aging archetypes, offering a rigorous, generative paradigm for simulating human molecular aging.
Paulo R. Ferreira, Lucas Coutinho Freitas, Laís dos Santos Gonçalves +4cs.LG q-bio.GN
NA methylation profiling has become a powerful approach for central nervous system (CNS) tumor classification, yet important challenges remain regarding cross-cohort transferability, methodological correctness, and robust multiclass evaluation. In this work, we propose a novel and methodologically rigorous machine-learning approach for methylation-based CNS tumor classification that combines Sparse Random Projection for dimensionality reduction with multinomial logistic regression for classification. We evaluate the proposed approach in the same general experimental setting established by a widely used reference classifier. On the 2,801-sample reference cohort, our method achieves a mean accuracy of 96\% under stratified 3-fold cross-validation. On the independent 1,104-sample clinical evaluation cohort, it reaches 86\% accuracy at the 91-class level and 93\% when predictions are evaluated at the methylation class family level. These results improve upon the corresponding state-of-the-art reference figures of 82\% class-level concordance and 88\% family-level concordance, yielding absolute gains of approximately 4 and 5 percentage points, respectively. This improvement is clinically relevant: in a diagnostic setting, a 5-point increase in correct tumor classification can directly affect cancer subtype assignment and, in turn, influence treatment selection and downstream clinical decision-making. Our results show that the proposed model, grounded in stronger methodological practice in machine learning, consistently outperforms the previous state of the art across evaluation settings and can materially improve the reliability of CNS tumor classification.