Joint Modeling of High-Dimensional Omics Data Through Latent Factors and Time-to-Event Data

seminar flyer

Daniel J. Schaid, Ph.D.

Enterprise Chair/Interim Chair
Professor of Biostatistics
Dept. of Quantitative Health Sciences

Mayo Clinic

High-dimensional omics data are increasingly used to identify biomarkers and biological mechanisms associated with clinical outcomes. However, typical studies measure thousands of molecular features in relatively small cohorts, limiting traditional analysis methods. There is a need for statistically efficient methods that jointly model high dimensional omics data and censored time-to-event outcomes to improve inference and biological interpretability.

We propose a joint latent factor model that simultaneously estimates low dimensional representations of high-dimensional omics data and their associations with censored time-to-event outcomes. The model incorporates a
bilinear matrix linear structure with elastic net penalization to achieve sparse and interpretable factor loadings, while incorporating informative annotations of the omics data. By jointly estimating latent structure and survival associations, the method improves statistical efficiency compared to two-stage approaches. Application to a metastatic breast cancer study (PROMISE; N=45) integrating 710 metabolites, 750 lipid s, and 1,473 proteins demonstrated that the joint model identified metabolite-derived latent factors significantly associated with time to disease progression, whereas marginal analyses of individual metabolites did not achieve multiple testing significance. Penalized loadings highlighted amino acid metabolites as the primary contributors to progression risk, illustrating the method’s ability to detect coordinated weak effects across biologically related features.