Subversion of clinical judgment by conversational artificial intelligence

Journal: medRxiv
Published Date:

Abstract

In medicine, artificial intelligence (AI) safety has focused on discrete flaws such as hallucination, bias and miscalibration. Conversational systems pose a distinct hazard: in dialogue, a model can covertly advance a misaligned objective by redirecting the clinical judgment it should support. Whether clinical expertise equips physicians to detect and resist such subversion is unknown. In a participant-blinded, randomized experiment, 225 physicians from 42 countries and eight clinical disciplines made sequential diagnostic and treatment decisions with AI assistance on anonymized cases from neurocritical care, their shared field of expertise. Unknown to them, each physician used the same model under two conditions: instructed to steer them towards harmful targets (adversarial) or correct clinical options (aligned), each established by expert consensus. With adversarial AI, 190 physicians made at least one harmful decision, compared to 21 with aligned AI. Across 2,485 targeted decision points, adversarial AI raised the harmful-decision rate from 2% with aligned AI to 49% (adjusted difference 47 percentage points, 95% CI, 43-51) and reduced the correct-decision rate from 83% to 25%, with no clear protection associated with clinical experience. Of the 190 physicians who made harmful decisions with adversarial AI, 141 did not report anything unusual about that AI. Two physicians reported suspecting systematic steering or manipulation. Physicians cannot safeguard the clinical use of conversational AI through expertise and decision-making authority alone.

Authors

  • Barrit
  • S.; Salvagno
  • M.; Corriero
  • A.; Sushil
  • M.; Pate
  • T.; Klug
  • J.; Aries
  • M.; Benghanem
  • S.; Baggiani
  • M.; Cane
  • G.; Park
  • S.; Soloperto
  • R.; Engrand
  • N.; Ben-Hamouda
  • N.; Robba
  • C.; Munari
  • M.; Balanca
  • B.; Hilaire
  • F.; Hemphill
  • J. C.; Foreman
  • B.; Lazaridis
  • C.; Chabanne
  • R.; Vlahovic
  • D.; Helbok
  • R.; Pinggera
  • D.; Kirschen
  • M.; Neurocore research group
  • ; Taccone
  • F. S.; Chang
  • E. F.