Paper List

Computational Neuroscience

Translating Measures onto Mechanisms: The Cognitive Relevance of Higher-Order Information

2025-12-02

This review addresses the core challenge of translating abstract higher-order information theory metrics (e.g., synergy, redundancy) into defensible, ...
Artificial Intelligence

Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs

2025-12-02

This paper addresses the critical gap in understanding whether LLMs spontaneously develop human-like Bayesian strategies for processing uncertain info...
Bioinformatics

Vessel Network Topology in Molecular Communication: Insights from Experiments and Theory

2025-12-02

This work addresses the critical lack of experimentally validated channel models for molecular communication within complex vessel networks, which is ...
Biophysics

Modulation of DNA rheology by a transcription factor that forms aging microgels

2025-12-02

This work addresses the fundamental question of how the transcription factor NANOG, essential for embryonic stem cell pluripotency, physically regulat...
Systems Biology

Imperfect molecular detection renormalizes apparent kinetic rates in stochastic gene regulatory networks

2025-12-02

This paper addresses the core challenge of distinguishing genuine stochastic dynamics of gene regulatory networks from artifacts introduced by imperfe...
Bioinformatics

PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer

2025-12-02

This paper addresses the dual challenge of achieving computational efficiency without sacrificing accuracy in whole-transcriptome single-cell represen...
Mathematical Biology

Beyond Bayesian Inference: The Correlation Integral Likelihood Framework and Gradient Flow Methods for Deterministic Sampling

2025-12-02

This paper addresses the core challenge of calibrating complex biological models (e.g., PDEs, agent-based models) with incomplete, noisy, or heterogen...
Bioinformatics

Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing

2025-12-02

This paper addresses the core challenge of detecting viral variants in wastewater sequencing data without reference genomes or labeled annotations, ov...

14 / 18

期刊: ArXiv Preprint

发布日期: 2026-03-07

BioinformaticsComputational Chemistry

MS2MetGAN: Latent-space adversarial training for metabolite–spectrum matching in MS/MS database search

University of Tennessee at Chattanooga | Middle Georgia State University

Yingfeng Wang, Meng Tsai, Alexzander Dwyer, Estelle Nuckels

30秒速读

IN SHORT: This paper addresses the critical bottleneck in metabolite identification: the generation of high-quality negative training samples that are structurally similar to true metabolites, which is essential for training robust machine learning classifiers.

核心创新

Methodology Introduces a novel latent-space approach where metabolite structures and MS/MS spectra are encoded into numerical vectors using autoencoders, transforming metabolite-spectrum matching into vector matching.
Methodology Proposes a GAN framework specifically designed to generate challenging decoy metabolite latent vectors conditioned on spectrum latent vectors, creating more informative negative training samples.
Methodology Demonstrates that adversarial training (GAN-9) significantly improves classifier stability, reducing standard deviation of accuracy across datasets from 0.3286 (GAN-0) to 0.1618 while increasing mean accuracy.

主要结论

MS2MetGAN achieves superior overall performance with mean accuracy of 76.33% against MetaCyc database and 79.35% against isomer decoys, outperforming 8 baseline tools including MIDAS (69.21%), SF-Matching (65.79%), and CSI:FingerID (49.66%).
The GAN training procedure improves performance stability across diverse test datasets, reducing standard deviation of accuracy from 0.3286 (GAN-0) to 0.1618 (GAN-9) for MetaCyc searches and from 0.3122 to 0.1663 for isomer decoy searches.
MS2MetGAN demonstrates strong transferability, outperforming baseline tools on 66.67%-100% of test datasets in pairwise comparisons, with particularly strong performance against isomer decoys where it beats all baselines on 77.78%-100% of datasets.

研究空白： Current metabolite identification tools are limited by the quality of negative training samples, typically constructed from mismatched compounds or spectra that are not structurally similar enough to true metabolites, leading to suboptimal classifier performance.

摘要: Database search is a widely used approach for identifying metabolites from tandem mass spectra (MS/MS). In this strategy, an experimental spectrum is matched against a user-specified database of candidate metabolites, and candidates are ranked such that true metabolite–spectrum matches receive the highest scores. Machine-learning methods have been widely incorporated into database-search–based identification tools and have substantially improved performance. To further improve identification accuracy, we propose a new framework for generating negative training samples. The framework first uses autoencoders to learn latent representations of metabolite structures and MS/MS spectra, thereby recasting metabolite–spectrum matching as matching between latent vectors. It then uses a GAN to generate latent vectors of decoy metabolites and constructs decoy metabolite–spectrum matches as negative samples for training. Experimental results show that our tool, MS2MetGAN, achieves better overall performance than existing metabolite identification methods.