Tutorials

Nonlinear multi-study sparse factor analysis

arXiv:2601.18128v2 Announce Type: replace-cross Abstract: High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studi

DGX agentpaper
tutorialsarxiv-cs-lg

arXiv:2601.18128v2 Announce Type: replace-cross Abstract: High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studies, one goal is to understand which underlying factors are common to all studies, and which factors are study-specific. As a particular example, we consider platelet gene expression data from patients in different disease groups. In this data, factors correspond to clusters of genes which are co-expressed; we may expect some clusters (or biological pathways) to be active for all diseases, while some clusters are only active for a specific disease. To learn these factors, we consider a nonlinear multi-study sparse factor model, which allows for both shared and study-specific factors. To fit this model, we propose a multi-study sparse variational autoencoder. The underlying model is sparse in that each observed feature (i.e. each dimension of the data) depends on a small subset of the latent factors. In the genomics example, this means each gene is active in only a few biological processes. We prove that the latent factor distributions are identifiable (up to element-wise transformations), and that the canonical shared and study-specific support structure is identifiable. Empirically, we demonstrate our method recovers meaningful factors in the platelet gene expression data.

Source: arXiv cs.LG | 2026-08-12

Loading related sources…