Combining Search, Social Media, and Traditional Data Sources to Improve Influenza Surveillance
Top Cited Papers
Open Access
- 29 October 2015
- journal article
- research article
- Published by Public Library of Science (PLoS) in PLoS Computational Biology
- Vol. 11 (10), e1004513
- https://doi.org/10.1371/journal.pcbi.1004513
Abstract
We present a machine learning-based methodology capable of providing real-time (“nowcast”) and forecast estimates of influenza activity in the US by leveraging data from multiple data sources including: Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system. Our main contribution consists of combining multiple influenza-like illnesses (ILI) activity estimates, generated independently with each data source, into a single prediction of ILI utilizing machine learning ensemble approaches. Our methodology exploits the information in each data source and produces accurate weekly ILI predictions for up to four weeks ahead of the release of CDC’s ILI reports. We evaluate the predictive ability of our ensemble approach during the 2013–2014 (retrospective) and 2014–2015 (live) flu seasons for each of the four weekly time horizons. Our ensemble approach demonstrates several advantages: (1) our ensemble method’s predictions outperform every prediction using each data source independently, (2) our methodology can produce predictions one week ahead of GFT’s real-time estimates with comparable accuracy, and (3) our two and three week forecast estimates have comparable accuracy to real-time predictions using an autoregressive model. Moreover, our results show that considerable insight is gained from incorporating disparate data streams, in the form of social media and crowd sourced data, into influenza predictions in all time horizons. The aggregated activity patterns of Internet users have enabled the detection and tracking of multiple population-wide events such as disease outbreaks, financial markets performance, and preferences in online movie selections. As a consequence, a collection of mathematical models aiming at monitoring and predicting these events in real-time have been proposed in the past decade. As we discover new methods and data sources suitable to track these events, it is not clear whether more information will lead to improved predictions. In the context of digital disease detection at the population level, we show that it is advantageous to combine the information from multiple flu activity predictors in the US instead of simply choosing the best performing flu predictor. Our findings suggest that the information from multiple data sources such as Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system, complement one another and produce the most accurate and robust set of flu predictions when combined optimally.This publication has 35 references indexed in Scilit:
- Using search queries for malaria surveillance, ThailandMalaria Journal, 2013
- Reassessing Google Flu Trends Data for Detection of Seasonal and Pandemic Influenza: A Comparative Epidemiological Study at Three Geographic ScalesPLoS Computational Biology, 2013
- Monitoring Influenza Epidemics in China with Search Query from BaiduPLOS ONE, 2013
- Optimizing Provider Recruitment for Influenza Surveillance NetworksPLoS Computational Biology, 2012
- Assessing Google Flu Trends Performance in the United States during the 2009 Influenza Virus A (H1N1) PandemicPLOS ONE, 2011
- A New Approach to Monitoring Dengue ActivityPLoS Neglected Tropical Diseases, 2011
- Using Web Search Query Data to Monitor Dengue Epidemics: A New Model for Neglected Tropical Disease SurveillancePLoS Neglected Tropical Diseases, 2011
- The Use of Twitter to Track Levels of Disease Activity and Public Concern in the U.S. during the Influenza A H1N1 PandemicPLOS ONE, 2011
- Detecting influenza epidemics using search engine query dataNature, 2009
- Monitoring the Impact of Influenza by Age: Emergency Department Fever and Respiratory Complaint Surveillance in New York CityPLoS Medicine, 2007