Researchers tested multimodal AI systems using annotated urban images from more than 200 cities. They examined both image object detection and subjective judgments of urban visual qualities, comparing results with and without geographic information such as coordinates, city, country or continent.

The systems showed different results across regions. After geographic information was added, perceptual scores for attributes including “Wealthy” and “Beautiful” fell by an average of 80% for African cities, while scores for Asia and North America rose by 26%. Among the models tested, GPT-4o showed the lowest perceptual bias.