Paper List
-
EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants
This paper addresses the core challenge of jointly predicting enzyme kinetic parameters (Kcat and Km) by modeling dynamic enzyme-substrate interaction...
-
Tissue stress measurements with Bayesian Inversion Stress Microscopy
This paper addresses the core challenge of measuring absolute, tissue-scale mechanical stress without making assumptions about tissue rheology, which ...
-
DeepFRI Demystified: Interpretability vs. Accuracy in AI Protein Function Prediction
This study addresses the critical gap between high predictive accuracy and biological interpretability in DeepFRI, revealing that the model often prio...
-
Hierarchical Molecular Language Models (HMLMs)
This paper addresses the core challenge of accurately modeling context-dependent signaling, pathway cross-talk, and temporal dynamics across multiple ...
-
Stability analysis of action potential generation using Markov models of voltage‑gated sodium channel isoforms
This work addresses the challenge of systematically characterizing how the high-dimensional parameter space of Markov models for different sodium chan...
-
Personalized optimization of pediatric HD-tDCS for dose consistency and target engagement
This paper addresses the critical limitation of one-size-fits-all HD-tDCS protocols in pediatric populations by developing a personalized optimization...
-
Consistent Synthetic Sequences Unlock Structural Diversity in Fully Atomistic De Novo Protein Design
This paper addresses the core pain point of low sequence-structure alignment in existing synthetic datasets (e.g., AFDB), which severely limits the pe...
-
Generative design and validation of therapeutic peptides for glioblastoma based on a potential target ATP5A
This paper addresses the critical bottleneck in therapeutic peptide design: how to efficiently optimize lead peptides with geometric constraints while...
Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
Computer Science, Old Dominion University | Biomedical Informatics, University of Arkansas for Medical Sciences
The 30-Second View
IN SHORT: This paper addresses the core challenge of generating accurate and clinically relevant patient notes from sparse inputs (ICD codes and basic demographics) by augmenting Chain-of-Thought prompting with semantic search and structured medical knowledge graphs.
Innovation (TL;DR)
- Methodology Proposes a novel hybrid prompting framework that integrates traditional Chain-of-Thought reasoning with semantic search results from a clinical corpus (CodiEsp dataset) to provide contextual examples.
- Methodology Introduces the infusion of a structured clinical ontology knowledge graph (built from SNOMED CT OWL expressions) directly into the LLM prompt to ground generation in formal medical relationships and constraints.
- Methodology/Biology Demonstrates the first systematic approach to reverse the common ICD code classification task, instead generating comprehensive clinical notes from ICD codes as primary input, evaluated on six distinct clinical cases.
Key conclusions
- The proposed CoT prompting with semantic search (using ICD codes as query) consistently outperformed the standard one-shot baseline across six clinical cases, as evidenced by lower cosine distance scores (e.g., Case C showed a clear leftward shift in KDE peak, indicating higher semantic similarity to ground truth).
- Incorporating a clinical knowledge graph (SNOMED CT OWL) into the prompt (CoT KG) provided structured medical relationships, enriching the generated notes with domain-specific terminology and logical constraints derived from formal ontologies.
- The hybrid approach (CoT Semantic Search + KG) leverages both in-context examples from similar cases and formal medical knowledge, offering a robust framework for improving the factual accuracy and clinical relevance of LLM-generated notes from coded inputs.
Abstract: In the past decade a surge in the amount of electronic health record (EHR) data in the United States, attributed to a favorable policy environment created by the Health Information Technology for Economic and Clinical Health (HITECH) Act of 2009 and the 21st Century Cures Act of 2016. Clinical notes for patients’ assessments, diagnoses, and treatments are captured in these EHRs in free-form text by physicians, who spend a considerable amount of time entering and editing them. Manually writing clinical notes takes a considerable amount of a doctor’s valuable time, increasing the patient’s waiting time and possibly delaying diagnoses. Large language models (LLMs) possess the ability to generate news articles that closely resemble human-written ones. We investigate the usage of Chain-of-Thought (CoT) prompt engineering to improve the LLM’s response in clinical note generation. In our prompts, we use as input International Classification of Diseases (ICD) codes and basic patient information. We investigate a strategy that combines the traditional CoT with semantic search results to improve the quality of generated clinical notes. Additionally, we infuse a knowledge graph (KG) built from clinical ontology to further enrich the domain-specific knowledge of generated clinical notes. We test our prompting technique on six clinical cases from the CodiEsp test dataset using GPT-4 and our results show that it outperformed the clinical notes generated by standard one-shot prompts.