Keyphrase Identification Using Minimal Labeled Data with Hierarchical Contexts and Transfer Learning

Journal: medRxiv
Published Date:

Abstract

Background: Interoperable clinical decision support system (CDSS) rules provide a pathway to interoperability, a well-recognized challenge in health information technology. Building an ontology facilitates creating interoperable CDSS rules, and identifying the keyphrases (KP) from the existing literature can be a first step for building an ontology. Ontology construction requires curation by human domain experts (HDE) traditionally, and modern natural language processing (NLP) techniques can be a critical complementary component nevertheless requires human proficiency, consensus, and contextual understanding for data labeling. Methods: We present a semi-supervised KP identification framework using a hierarchical-attention BiLSTM-CRF (Hier-Attn-BiLSTM-CRF) with word-, sentence-, and document-level attention. A domain-adapted sciSpacy model was used to generate synthetic labels for bootstrap training, followed by fine-tuning with minimal HDE-labeled data. We then evaluated robustness through component ablation, comparison with fine-tuned biomedical transformers (BioBERT, SciBERT, PubMedBERT) with bootstrap confidence intervals, train-split sensitivity, multi-seed variance analysis, error analysis, and benchmarking against public corpora (KPBioMed, PubMedAKE). Results: The Hier-Attn-BiLSTM-CRF is competitive with fine-tuned transformer baselines on the HDE labeled dataset (GS42: ~44 vs. ~46 F1; GS91: 61.3 vs. ~54 F1), through an explicit hierarchical inductive bias over word-, sentence-, and document-level representations, complementing the implicit contextual modeling of the transformer. Controlled ablation identifies gold-standard fine-tuning as the dominant performance lever. Mixing HDE and synthetic labels in 2:100-4:100 ratios improved performance without exhausting the human-labeled set too quickly. Models trained on sparser public corpora transferred poorly to CDSS across all architectures, underscoring the value of in-domain synthetic labels. Conclusions: This feasibility study demonstrates a practical, resource-efficient framework for CDSS KP identification under limited HDE annotation. The contribution lies in integrating established components: domain-adapted synthetic labels, hierarchical attention, and minimal gold-standard fine-tuning, for the CDSS sub-domain. A full downstream evaluation of the role of the NLP pipeline for CDSS ontology construction is the primary next step.

Authors

  • Goli
  • R.; Komatineni
  • K.; Alluri
  • S.; Hubig
  • N.; Min
  • H.; Gong
  • Y.; Sitting
  • D.; Rennert
  • L.; Robinson
  • D.; Biondich
  • P.; Wright
  • A.; Nohr
  • C.; Law
  • T.; Faxvaag
  • A.; Weaver
  • A.; Gimbel
  • R.; Jing
  • X.