Evaluating performance bias in face-to-BMI vision transformer models across diverse human populations
Journal:
bioRxiv
Published Date:
Sep 4, 2026
Abstract
Computer vision models that estimate body mass index (BMI) from facial features offer a non-invasive, low-cost alternative to physical measurement, with uses in telemedicine, emergency care where a scale or measuring tools arent available, automated self-monitoring, and large-scale epidemiological research. Most of these models, however, are trained on government records, social media images, and celebrity photographs, sources that introduce dataset biases and fail to represent the general public. This study tests how well a face-to-BMI machine learning model generalizes across populations, specifically how morphological diversity and population-specific training data affect cross-cultural accuracy. We trained and evaluated Vision Transformer (ViT-H/14) models on paired BMI measurements and facial photographs from four Indigenous populations: the Orang Asli of Malaysia, the Ju/hoansi of Southern Africa, the Sama residing in the Philippines, and the Tsimane of Bolivia. To evaluate how training data composition affects predictions, we compared four training strategies, from single-population models (focal models) to models trained on the full combined global dataset (global models). In-distribution training always produced the best performance. Models exposed to a target populations morphology, whether focal or global, consistently predicted BMI most accurately for that population. But when a target population differed from the training sample, adding more cross-cultural variation to training improved out-of-distribution predictions. Therefore, training on a populations own data works best when that data exists, and training on data spanning a wide range of human morphology is the strongest fallback when it doesnt. These findings suggest that while target population training data produces the most accurate results, training on datasets that capture global morphological variation substantially improves performance in unrepresented populations. Broader diversity in training data is essential for developing machine learning health tools that generalize reliably across human populations.