Precision Evidence Bench: Assessing PICO alignment and faithful reporting in clinical AI

Journal: medRxiv
Published Date:

Abstract

Clinicians routinely make decisions for patients whose comorbidities, prior therapies, laboratory abnormalities, or demographic characteristics are not reflected in the supporting evidence, even though these factors can influence treatment selection. Given this gap in the evidence base, we sought to evaluate how AI models respond to clinical questions in which patient characteristics or the treatment comparison and outcome of interest may affect the applicability of available evidence. We developed and publicly released Precision Evidence Bench, a benchmark of 209 synthetic clinical questions with specified population, intervention, comparator, and outcome (PICO) elements. Each answer is assessed using a five item rubric that evaluates whether a single cited source matches the PICO and whether the answer faithfully reports its sources. The item verdicts determine one of five ratings, ranging from Precision match to Unsupported claim, without requiring a reference answer. We evaluated four frontier large language models (LLMs) using web search across all 209 questions, then repeated the evaluation with additional access to a corpus of precision evidence-based findings (pEBFs) designed to address these evidence gaps, using two corpus sizes. For every model, Precision match rates increased with the amount of precision evidence available, from 13 to 18% with web search alone to 17 to 28% with 100 million pEBFs and 41 to 56% with 500 million pEBFs. These findings suggest that access to matching evidence is an important determinant of LLM performance on precision medicine questions.

Authors

  • Mukerji
  • A.; Sanghavi
  • N.; Lauritzen
  • T.; Muralidharan
  • J.; Gombar
  • S.