How Biased is the Population of Facebook Users? Comparing the Demographics of Facebook Users with Census Data to Generate Correction Factors
- 6 July 2020
- conference paper
- conference paper
- Published by Association for Computing Machinery (ACM)
Abstract
Censuses and representative sampling surveys around the world are key sources of data to guide government investments and public policies. However, these sources are very expensive to obtain and are collected relatively infrequently. Over the last decade, there has been growing interest in the use of data from social media to complement more traditional data sources. However, social media users are not representative of the general population. Thus, analyses based on social media data require statistical adjustments, like post-stratification, in order to remove the bias and make solid statistical claims. These adjustments are possible only when we have information about the frequency of demographic groups using social media. These data, when compared with official statistics, enable researchers to produce appropriate statistical correction factors. In this paper, we leverage the Facebook advertising platform to compile the equivalent of an aggregate-level census of Facebook users. Our compilation includes the population distribution for seven demographic attributes such as gender, political leaning, and educational attainment at different geographic levels for the U.S. (country, state, and city). By comparing the Facebook counts with official reports provided by the U.S. Census and Gallup, we found very high correlations, especially for political leaning and race. We also identified instances where official statistics may be underestimating population counts as in the case of immigration. We use the information collected to calculate bias correction factors for all computed attributes in order to evaluate the extent to which different demographic groups are more or less represented on Facebook, and to derive the actual distributions for specific audiences of interest. We provide the first comprehensive analysis for assessing biases in Facebook users across several dimensions. This information can be used to generate bias-adjusted population estimates and demographic counts in a timely way and at fine geographic granularity in between data releases of official statistics.Keywords
This publication has 18 references indexed in Scilit:
- The Impact of Hurricane Maria on Out‐migration from Puerto Rico: Evidence from Facebook DataPopulation and Development Review, 2019
- A Large-Scale Analysis of Facebook’s User-Base and User Engagement GrowthIEEE Access, 2018
- Promises and Pitfalls of Using Digital Traces for Demographic ResearchDemography, 2018
- Studying Migrant Assimilation Through Facebook InterestsPublished by Springer Science and Business Media LLC ,2018
- Using Facebook ad data to track the global digital gender gapWorld Development, 2018
- Analyzing gender inequality through large-scale Facebook advertising dataProceedings of the National Academy of Sciences of the United States of America, 2018
- Online Health Monitoring using Facebook Advertisement Audience Estimates in the United States: Evaluation StudyJMIR Public Health and Surveillance, 2018
- Using Facebook Ads Audiences for Global Lifestyle Disease SurveillancePublished by Association for Computing Machinery (ACM) ,2017
- Predicting political preference of Twitter usersPublished by Association for Computing Machinery (ACM) ,2013
- Computing political preference among twitter followersPublished by Association for Computing Machinery (ACM) ,2011