A 'Silent Trial' Assessing the Accuracy of Large Language Models for Assisting Community Health Workers in Low-Resource Settings

Journal: medRxiv
Published Date:

Abstract

Community health workers (CHWs) in low-resource settings deliver variable-quality care. This study used OpenAI's o3 and Google's Gemini Flash 2.5 to evaluate whether large language models (LLMs) 'listening' to CHW-patient interactions could generate accurate referral decisions. Across 150 participating Rwandan CHWs, 429 encounters were recorded (in Kinyarwanda) and then processed by LLMs. CHWs demonstrated high referral accuracy (97.9% [95% CI: 96.1%-98.9%]), and OpenAI's o3 performed similarly to CHWs while Gemini 2.5-Flash showed low accuracy (47.3% [95% CI: 42.6%-52.1%]). Assessment of LLM-generated differential diagnoses and management plan quality showed superior performance from o3 compared with Gemini, though both models missed important conditions. In conclusion, the choice of LLM appears to be a critical design decision. Moreover, the high baseline performance of Rwandan CHWs suggests that LLMs are likely to have a limited impact in the current context but could be useful in less well-established CHW programmes. Trial Registration: PACTR202504601308784.

Authors

  • Shimelash
  • N.; Rutunda
  • S.; Menon
  • V.; Emmanual-Fabula
  • M.; Uwimbabazi
  • A.; Rugege
  • C.; Nshimiyimana
  • C.; Rwema
  • I.; Kandekwe
  • M.; Berhe
  • D. F. D.; Wong
  • R.; Remera
  • E.; Hezagira
  • E.; Gill
  • J.; Archer
  • L.; Riley
  • R. D.; Denniston
  • A. K.; Liu
  • X.; Mateen
  • B.