Evaluating in-context learning with prompting regimes for transformer-based synthetic health tabular data generation

Journal: bioRxiv
Published Date:

Abstract

Tabular synthetic data generation (SDG) can facilitate research and development in the health domain, as access to health data is restricted. To accommodate for its structural complexity, such as multimodal distributions and fragmentation, transformer models provides an alternative to conventional SDG methods since they are better equipped to infer global spatial coherence due to textual context of tabular records. However, methods for optimising their adaption to tabular SDG are still largely unexplored. We contribute to the understanding of transformer-based SDG by evaluating fidelity and privacy of three prompting regimes, zero shot, one shot and few shot training, for tabular health SDG using small transformer models. We evaluated llama 3.2 (1B and 3B parameters) and Phi (1.3B and 2.7B parameters), on a real clinical dataset. Our results showed that zero shot learning was less prone to hallucinations and performed overall better for fidelity, but at the loss of privacy in comparison to the other regimes. Our results showed no differences between model sizes and families. Future studies should extend the analysis to other models and datasets, and explore strategies to reduce hallucinations.

Authors

  • Bertgren
  • A.; Ohberg
  • F.; Soda
  • P.; Naslund
  • U.; Wennberg
  • P.; Gronlund
  • C.

Categories