Society & Economicsarticle2026-08-04

A tale of ten cities: data bias in human mobility is pervasive and highly location specific

Open access0 citations

Abstract

Abstract Large-scale human mobility datasets play increasingly critical roles in many algorithmic systems, business processes, and policy decisions. Unfortunately, there has been little focus on understanding bias and other fundamental shortcomings of these datasets and how they impact downstream analyses and prediction tasks. In this work, we study ‘data production’, quantifying not only whether individuals are represented in big digital datasets, but also how they are represented in terms of how much data they produce. We study one GPS mobility dataset which is collected from anonymized smartphones for ten major US cities and find that data points can be more unequally distributed between users than wealth. We build models to predict the number of data points we can expect to be produced by the composition of demographic groups living in census tracts, and find strong effects of wealth, ethnicity, and education on data production. While we find that bias is a ubiquitous phenomenon, occurring in all ten cities, we further find that each city suffers from its own manifestation of it, and that location-specific models are required to model bias for each city. This work raises serious questions about general approaches to debias human mobility data and urges further research.

// Source

View paper (DOI)Open access versionOpenAlexEPJ Data SciencePublished 2026-08-04

Authors: Katinka den Nijs, Elisa Omodei, Vedran Sekara

Institutions: Pioneer (United States), IT University of Copenhagen, Central European University