Paper List

Bioinformatics

SpikGPT: A High-Accuracy and Interpretable Spiking Attention Framework for Single-Cell Annotation

2025-12-02

This paper addresses the core challenge of robust single-cell annotation across heterogeneous datasets with batch effects and the critical need to ide...
Bioinformatics

Unlocking hidden biomolecular conformational landscapes in diffusion models at inference time

2025-12-02

This paper addresses the core challenge of efficiently and accurately sampling the conformational landscape of biomolecules from diffusion-based struc...
Computational Neuroscience

Personalized optimization of pediatric HD-tDCS for dose consistency and target engagement

2025-12-01

This paper addresses the critical limitation of one-size-fits-all HD-tDCS protocols in pediatric populations by developing a personalized optimization...
Computational Biophysics

Realistic Transition Paths for Large Biomolecular Systems: A Langevin Bridge Approach

2025-12-01

This paper addresses the core challenge of generating physically realistic and computationally efficient transition paths between distinct protein con...
Bioinformatics

Consistent Synthetic Sequences Unlock Structural Diversity in Fully Atomistic De Novo Protein Design

2025-12-01

This paper addresses the core pain point of low sequence-structure alignment in existing synthetic datasets (e.g., AFDB), which severely limits the pe...
Bioinformatics

MoRSAIK: Sequence Motif Reactor Simulation, Analysis and Inference Kit in Python

2025-12-01

This work addresses the computational bottleneck in simulating prebiotic RNA reactor dynamics by developing a Python package that tracks sequence moti...
Bioinformatics

On the Approximation of Phylogenetic Distance Functions by Artificial Neural Networks

2025-12-01

This paper addresses the core challenge of developing computationally efficient and scalable neural network architectures that can learn accurate phyl...
Bioinformatics

EcoCast: A Spatio-Temporal Model for Continual Biodiversity and Climate Risk Forecasting

2025-12-01

This paper addresses the critical bottleneck in conservation: the lack of timely, high-resolution, near-term forecasts of species distribution shifts ...

15 / 18

期刊: ArXiv Preprint

发布日期: 2026-03

BioinformaticsData Integration

Open Biomedical Knowledge Graphs at Scale: Construction, Federation, and AI Agent Access with Samyama Graph Database

VaidhyaMegha Private Limited, India

Madhulatha Mandarapu, Sandeep Kunkunuru

30秒速读

IN SHORT: This paper addresses the core pain point of fragmented biomedical data by constructing and federating large-scale, open knowledge graphs to enable seamless cross-domain queries and natural language access via AI agents.

核心创新

Methodology A reproducible ETL pattern for constructing large-scale biomedical KGs from heterogeneous public sources, featuring cross-source deduplication, batch loading, and portable snapshot export.
Methodology Demonstration of cross-KG federation via property-based joins, enabling queries that traverse multiple independent knowledge graphs (e.g., from clinical trials to biological pathways).
Methodology Schema-driven MCP server generation that automatically exposes typed tools for LLM agents, achieving 98% accuracy on a new BiomedQA benchmark, significantly outperforming text-to-Cypher (0%) and standalone GPT-4o (75%).

主要结论

The federated graph (7.9M nodes, 28M edges) loads in approximately 3 minutes on commodity hardware (AWS g4dn.4xlarge, 62 GB RAM), with cross-KG queries completing in 80 ms–4 s.
Schema-driven MCP tools achieve 98% accuracy (39/40) on the BiomedQA benchmark, dramatically outperforming text-to-Cypher (0%) and standalone GPT-4o (75%).
A Rust native loader constructs the Drug Interactions KG (32,726 nodes, 191,970 edges) in under 1 second, demonstrating orders-of-magnitude performance improvement over Python HTTP-based ETL.

研究空白： Current biomedical knowledge integration efforts are limited by static datasets, lack of update pipelines, infrastructure complexity, and the inability to perform seamless cross-database queries, creating a bottleneck for translational research.

摘要: Biomedical knowledge is fragmented across siloed databases—Reactome for pathways, STRING for protein interactions, Gene Ontology for functional annotations, ClinicalTrials.gov for study registries, DrugBank for drug vocabularies, DGIdb for drug–gene interactions, SIDER for side effects, and dozens more. Researchers routinely download flat files from each source and write bespoke scripts to cross-reference them, a process that is slow, error-prone, and not reproducible. We present three open-source biomedical knowledge graphs—Pathways KG (118,686 nodes, 834,785 edges from 5 sources), Clinical Trials KG (7,774,446 nodes, 26,973,997 edges from 5 sources), and Drug Interactions KG (32,726 nodes, 191,970 edges from 3 sources)—built on Samyama, a high-performance graph database written in Rust. Our contributions are threefold. First, we describe a reproducible ETL pattern for constructing large-scale KGs from heterogeneous public data sources, with cross-source deduplication, batch loading (both Python Cypher and Rust native loaders), and portable snapshot export. Second, we demonstrate cross-KG federation: loading all three snapshots into a single graph tenant enables property-based joins across datasets, answering questions like “For drugs indicated for diabetes, what are their gene targets and which biological pathways do those targets participate in?”—a query that no single KG can answer alone. Third, we introduce schema-driven MCP server generation: each KG automatically exposes typed tools for LLM agents via the Model Context Protocol, enabling natural-language access to graph queries. We evaluate domain-specific MCP tools against text-to-Cypher and standalone GPT-4o on a new BiomedQA benchmark (40 pharmacology questions), achieving 98% accuracy vs. 0% for text-to-Cypher and 75% for standalone GPT-4o. All data sources are open-license (CC BY 4.0, CC0, OBO, public domain). Snapshots, ETL code, and MCP configurations are publicly available. The combined federated graph (7.9M nodes, 28M edges) loads in approximately 3 minutes from portable snapshots on commodity cloud hardware (AWS g4dn.4xlarge, 62 GB RAM), and cross-KG queries complete in 80 ms–4 s.

代码