Paper List
-
MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
This paper addresses the critical gap in healthcare AI systems that lack contextual reasoning, long-term state management, and verifiable workflows by...
-
Model Gateway: Model Management Platform for Model-Driven Drug Discovery
This paper addresses the critical bottleneck of fragmented, ad-hoc model management in pharmaceutical research by providing a centralized, scalable ML...
-
Tree Thinking in the Genomic Era: Unifying Models Across Cells, Populations, and Species
This paper addresses the fragmentation of tree-based inference methods across biological scales by identifying shared algorithmic principles and stati...
-
SSDLabeler: Realistic semi-synthetic data generation for multi-label artifact classification in EEG
This paper addresses the core challenge of training robust multi-label EEG artifact classifiers by overcoming the scarcity and limited diversity of ma...
-
Decoding Selective Auditory Attention to Musical Elements in Ecologically Valid Music Listening
This paper addresses the core challenge of objectively quantifying listeners' selective attention to specific musical components (e.g., vocals, drums,...
-
Physics-Guided Surrogate Modeling for Machine Learning–Driven DLD Design Optimization
This paper addresses the core bottleneck of translating microfluidic DLD devices from research prototypes to clinical applications by replacing weeks-...
-
Mechanistic Interpretability of Antibody Language Models Using SAEs
This work addresses the core challenge of achieving both interpretability and controllable generation in domain-specific protein language models, spec...
-
Fluctuating Environments Favor Extreme Dormancy Strategies and Penalize Intermediate Ones
This paper addresses the core challenge of determining how organisms should tune dormancy duration to match the temporal autocorrelation of their envi...
Probabilistic Joint and Individual Variation Explained (ProJIVE) for Data Integration
Department of Biostatistics and Bioinformatics, Rollins School of Public Health, Emory University | Department of Radiology and Imaging Sciences, Emory University School of Medicine
30秒速读
IN SHORT: This paper addresses the core challenge of accurately decomposing shared (joint) and dataset-specific (individual) sources of variation in multi-modal datasets, where existing methods often lack a formal statistical model, leading to potential inaccuracies and interpretability issues.
核心创新
- Methodology Introduces ProJIVE, a novel probabilistic model that extends Probabilistic PCA (pPCA) to the JIVE framework, formally modeling joint and individual subject scores as random effects.
- Methodology Develops a unified Expectation-Maximization (EM) algorithm for maximum likelihood estimation, simultaneously inferring all model parameters (loadings, scores, noise variances), unlike multi-step decomposition approaches.
- Biology Successfully applies the model to integrate brain morphometry and cognitive data from the ADNI cohort, demonstrating that the extracted joint scores strongly correlate with established but expensive Alzheimer's disease biomarkers (e.g., amyloid PET, FDG-PET, ApoE4 status).
主要结论
- ProJIVE's maximum likelihood estimation via EM achieved greater accuracy in estimating latent scores and variable loadings compared to R.JIVE, AJIVE, and GIPCA across various simulation settings, including non-Gaussian data.
- In the ADNI application, the joint subject scores derived from brain morphometry and cognition data showed strong statistical associations with key Alzheimer's disease variables, validating the biological relevance of the extracted shared variation.
- The model provides a formal statistical framework where quantities like joint subject scores (potential prodromes) and variable loadings (drivers of variation) are directly modeled, enhancing interpretability over algorithmic decompositions.
摘要: Collecting multiple types of data on the same set of subjects is common in modern scientific applications including genomics, metabolomics, and neuroimaging. Joint and Individual Variation Explained (JIVE) seeks a low-rank approximation of the joint variation between two or more sets of features captured on common subjects and isolates this variation from that unique to each set of features. We develop an expectation-maximization (EM) algorithm to estimate a probabilistic model for the JIVE framework. The model extends probabilistic PCA to multiple datasets. Our maximum likelihood approach simultaneously estimates joint and individual components, which can lead to greater accuracy compared to other methods. We apply ProJIVE to measures of brain morphometry and cognition in Alzheimer’s disease. ProJIVE learns biologically meaningful sources of variation, and the joint morphometry and cognition subject scores are strongly related to more expensive existing biomarkers. Data used in preparation of this article were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database. Code to reproduce the analysis is available at https://github.com/thebrisklab/ProJIVE. Supplementary materials for this article are available online.