A global test found that location information changed both what AI systems detected and how they judged urban scenes.
Researchers tested multimodal AI systems using annotated urban images from more than 200 cities. They examined both image object detection and subjective judgments of urban visual qualities, comparing results with and without geographic information such as coordinates, city, country or continent.
The systems showed different results across regions. After geographic information was added, perceptual scores for attributes including “Wealthy” and “Beautiful” fell by an average of 80% for African cities, while scores for Asia and North America rose by 26%. Among the models tested, GPT-4o showed the lowest perceptual bias.
How location changed AI judgments
The study found significant geographic differences in both object detection accuracy and judgments of urban images. Adding geographic information produced broadly similar directional changes whether the cue was a coordinate, city, country or continent; combining several geographic cues produced the strongest changes. After geo-referencing, scores for attributes such as “Wealthy” and “Beautiful” dropped by an average of 80% for African cities, while regions including Asia and North America showed score increases of 26%. GPT-4o had the lowest perceptual bias among the state-of-the-art models tested. The researchers report that the patterns resembled imbalances in training data, but the abstract does not establish that training data caused them.
What the test can show
The findings come from a standardized benchmark of annotated imagery from more than 200 cities, tested through a geographic red-teaming framework. The researchers compared model outputs across geographic regions and under different levels of location information. The abstract does not specify the full list of cities, images or models, the exact object-detection results, or how well the benchmark represents all urban environments. The tests show geographic disparities in model outputs, but they do not by themselves determine why those disparities occur or how the systems would perform in every real-world urban application.