PGxRAG: A Retrieval Augmented Generation supported Pharmacogenomics Assistant

Journal: medRxiv
Published Date:

Abstract

Pharmacogenomics enables personalized medicine by predicting individual drug responses based on genetic makeup, but complex guideline retrieval remains challenging for clinicians, particularly in resource-limited settings. While large language models (LLMs) show promise for numerous healthcare applications, their performance on domain-specific pharmacogenomics queries without expert knowledge integration remains limited. We evaluated whether Retrieval-Augmented Generation (RAG) enhancement improves LLM accuracy for pharmacogenomics applications compared to native model performance. We conducted comparative evaluation of four LLMs with and without RAG enhancement, constructing a knowledge base from PharmGKB (now ClinPGx), CPIC, Dutch Pharmacogenetics Working Group (KNMP), and FDA guidelines containing 2,617 embedded document chunks. We developed 225 multiple-choice questions representing patient and healthcare provider perspectives, then systematically evaluated hyperparameter combinations testing different temperatures, embedding dimensions, retrieval methods, and k-values with different sample sizes. RAG-enhanced models consistently outperformed native LLMs, with optimal configuration (GPT-4o) achieving 95.1% accuracy compared to 89.8% for the same native model. The RAG approach significantly enhances LLM performance in pharmacogenomics applications, providing a scalable solution for making complex pharmacogenomic guidelines accessible to healthcare providers while maintaining high clinical decision support accuracy.

Authors

  • Dhanush Borishetty; Peter Banda; Nikhilesh Andhi; Aakash Desai; Bharath Ram Uppili; Naga Mithil Samudrala; Shravani Shriya Palanki; Dheeraj Reddy Bobbili; Gayatri Rangarajan Iyer