PGxRAG: A Retrieval Augmented Generation supported Pharmacogenomics Assistant
Journal:
medRxiv
Published Date:
Jan 1, 2025
Abstract
Pharmacogenomics enables personalized medicine by predicting individual drug responses based on genetic makeup, but complex guideline retrieval remains challenging for clinicians, particularly in resource-limited settings. While large language models (LLMs) show promise for numerous healthcare applications, their performance on domain-specific pharmacogenomics queries without expert knowledge integration remains limited. We evaluated whether Retrieval-Augmented Generation (RAG) enhancement improves LLM accuracy for pharmacogenomics applications compared to native model performance. We conducted comparative evaluation of four LLMs with and without RAG enhancement, constructing a knowledge base from PharmGKB (now ClinPGx), CPIC, Dutch Pharmacogenetics Working Group (KNMP), and FDA guidelines containing 2,617 embedded document chunks. We developed 225 multiple-choice questions representing patient and healthcare provider perspectives, then systematically evaluated hyperparameter combinations testing different temperatures, embedding dimensions, retrieval methods, and k-values with different sample sizes. RAG-enhanced models consistently outperformed native LLMs, with optimal configuration (GPT-4o) achieving 95.1% accuracy compared to 89.8% for the same native model. The RAG approach significantly enhances LLM performance in pharmacogenomics applications, providing a scalable solution for making complex pharmacogenomic guidelines accessible to healthcare providers while maintaining high clinical decision support accuracy.