Paper List
-
Translating Measures onto Mechanisms: The Cognitive Relevance of Higher-Order Information
This review addresses the core challenge of translating abstract higher-order information theory metrics (e.g., synergy, redundancy) into defensible, ...
-
Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs
This paper addresses the critical gap in understanding whether LLMs spontaneously develop human-like Bayesian strategies for processing uncertain info...
-
Vessel Network Topology in Molecular Communication: Insights from Experiments and Theory
This work addresses the critical lack of experimentally validated channel models for molecular communication within complex vessel networks, which is ...
-
Modulation of DNA rheology by a transcription factor that forms aging microgels
This work addresses the fundamental question of how the transcription factor NANOG, essential for embryonic stem cell pluripotency, physically regulat...
-
Imperfect molecular detection renormalizes apparent kinetic rates in stochastic gene regulatory networks
This paper addresses the core challenge of distinguishing genuine stochastic dynamics of gene regulatory networks from artifacts introduced by imperfe...
-
PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
This paper addresses the dual challenge of achieving computational efficiency without sacrificing accuracy in whole-transcriptome single-cell represen...
-
Beyond Bayesian Inference: The Correlation Integral Likelihood Framework and Gradient Flow Methods for Deterministic Sampling
This paper addresses the core challenge of calibrating complex biological models (e.g., PDEs, agent-based models) with incomplete, noisy, or heterogen...
-
Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing
This paper addresses the core challenge of detecting viral variants in wastewater sequencing data without reference genomes or labeled annotations, ov...
SNPgen: Phenotype-Supervised Genotype Representation and Synthetic Data Generation via Latent Diffusion
DEIB, Politecnico di Milano | Health Data Science Centre, Human Technopole | Genomics Research Centre, Human Technopole | MOX - Department of Mathematics, Politecnico di Milano | Department of Public Health and Primary Care, University of Cambridge
30秒速读
IN SHORT: This paper addresses the core challenge of generating privacy-preserving synthetic genotype data that maintains both statistical fidelity and downstream predictive utility for supervised tasks like polygenic risk scoring.
核心创新
- Methodology Introduces a two-stage conditional latent diffusion framework combining GWAS-guided variant selection (1,024–2,048 SNPs) with VAE compression and phenotype-conditioned generation via classifier-free guidance.
- Methodology Implements phenotype-supervised generation rather than unconditional sampling, producing synthetic genotypes directly usable for downstream disease prediction tasks without additional phenotype mechanisms.
- Biology Demonstrates that GWAS-guided selection of trait-associated SNPs preserves predictive performance comparable to genome-wide methods while using 2–6× fewer variants, offering a favorable computational trade-off.
主要结论
- Models trained on synthetic data matched real-data predictive performance across four complex diseases (CAD, BC, T1D, T2D) in TSTR protocols, with synthetic XGBoost achieving AUCs of 0.587±0.019 for T2D and 0.594±0.011 for CAD, closely matching real-data performance.
- Privacy analysis showed zero identical matches, near-random membership inference (AUC ≈ 0.50), preserved LD structure, and high allele frequency correlation (r≥0.95) with source data, confirming strong privacy guarantees.
- In controlled simulations with known causal effects, synthetic data showed strong agreement with real-data effect estimates (Pearson r=0.835), exceeding VAE-reconstructed data (r=0.726), demonstrating faithful recovery of genetic association structures.
摘要: Motivation: Polygenic risk scores and other genomic analyses require large individual-level genotype datasets, yet strict data access restrictions impede sharing. Synthetic genotype generation offers a privacy-preserving alternative, but most existing methods operate unconditionally—producing samples without phenotype alignment—or rely on unsupervised compression, creating a gap between statistical fidelity and downstream task utility. Results: We present SNPgen, a two-stage conditional latent diffusion framework for generating phenotype-supervised synthetic genotypes. SNPgen combines GWAS-guided variant selection (1,024–2,048 trait-associated SNPs) with a variational autoencoder for genotype compression and a latent diffusion model conditioned on binary disease labels via classifier-free guidance. Evaluated on 458,724 UK Biobank individuals across four complex diseases (coronary artery disease, breast cancer, type 1 and type 2 diabetes), models trained on synthetic data matched real-data predictive performance in a train-on-synthetic, test-on-real protocol, approaching genome-wide PRS methods that use 2–6× more variants. Privacy analysis confirmed zero identical matches, near-random membership inference (AUC ≈ 0.50), preserved linkage disequilibrium structure, and high allele frequency correlation (r≥0.95) with source data. A controlled simulation with known causal effects verified faithful recovery of the imposed genetic association structure. Availability and implementation: Code available at https://github.com/ht-diva/SNPgen. Contact: andrea.lampis@polimi.it Supplementary information: Supplementary data are available in the Appendix.