Influenza-Like Illness Forecasting Using Multisource Data: Comparative Deep Learning Study.

Journal: JMIR medical informatics
Published Date:

Abstract

BACKGROUND: Accurate forecasting of influenza-like illness (ILI) is crucial for public health. Integrating novel digital data streams (eg, internet searches and human mobility) with traditional surveillance can improve accuracy, but optimal modeling frameworks are underexplored. OBJECTIVE: This study aimed to develop and compare multisource data-driven models for forecasting ILI incidence trends. METHODS: Weekly ILI incidence data and multisource variables for Hubei Province from 2020 to 2023 were collected. Multisource predictors included mean temperature, relative humidity, air quality index, a synthesized Baidu Search Index for influenza-related queries, the Baidu Migration Scale Index for population mobility, and the Oxford Stringency Index (SI) for nonpharmaceutical interventions. Predictive models evaluated were Seasonal Autoregressive Integrated Moving Average (SARIMA), long short-term memory (LSTM), a hybrid convolutional neural network-long short-term memory (CNN-LSTM), Transformer, and random forest. Models were constructed using an 85:15 training-test split and optimized via grid search with 5-fold cross-validation. Performance was assessed using mean absolute error, root-mean-squared error (RMSE), mean absolute percentage error, and coefficient of determination (R²). RESULTS: All models incorporating multisource data substantially outperformed univariate time-series benchmarks. Among univariate models, CNN-LSTM achieved superior performance (R²=0.7623) over SARIMA (R²=-0.2645). In multisource configurations, the LSTM model demonstrated the strongest predictive capability and feature integration, attaining an optimal R² of 0.8350 when combining environmental, mobility (Baidu Migration Scale), and internet search (synthetic Baidu index [SBI]) data. The inclusion of the Oxford SI consistently reduced prediction error across all models, with mean absolute percentage error decreasing by up to 84.87% in the LSTM model. The Baidu Search Index emerged as the most influential single external predictor, notably enhancing model fit. In contrast, model performance varied with feature composition: the Transformer model excelled with full feature sets, while CNN-LSTM performed best primarily with SBI integration. Over a 26-week prospective projection, the optimal LSTM model forecasted a "rapid decline-gradual decline-stabilization" trend, indicating a return to baseline ILI activity, and demonstrated robust external validation performance (R2=0.7707). CONCLUSIONS: Deep learning models, particularly LSTM, can effectively leverage heterogeneous digital data to improve ILI forecasting. Internet search behavior and policy stringency are critical external predictors. The proposed model demonstrated good external validation performance, and the findings may be applicable to other temperate regions in central China exhibiting similar epidemic characteristics.

Authors

Keywords

No keywords available for this article.