Researchers evaluated ResNet50, MobileNetV3 and the Data-efficient Image Transformer using pretrained weights and two facial-image datasets: FairFace, designed to provide balanced demographic groups, and the less controlled UTKFace dataset. They tested how well the models’ facial features supported image retrieval and demographic classification.
Face-recognition AI can score well overall while treating groups unevenly
Tests of three image models found that strong average performance did not remove disparities among race, gender and combined groups.

Uneven results across groups
A simple linear classifier consistently separated demographic information more effectively than one-nearest-neighbour retrieval on both datasets, showing that the pretrained facial representations contained demographic information that could be extracted with a basic classifier. ResNet50 had the highest overall classification performance, while MobileNetV3 produced more consistent subgroup results in several settings. Race–gender intersectional analysis found persistent demographic disparities even when overall accuracy improved. The study also evaluated fairness gaps, variation between groups and statistical significance, but the abstract does not report the numerical results of those tests.
Why average accuracy is not enough
Average accuracy can hide poorer performance for particular demographic groups. The findings support evaluating facial-recognition systems across individual and combined demographic groups rather than relying only on a single overall score.
What the tests show
The evidence comes from tests of three ImageNet-pretrained image models on the FairFace and UTKFace datasets, using retrieval and demographic-classification tasks. The study provides a framework for comparing representation quality and fairness, but results from these two datasets and these three models do not establish how all facial-recognition systems would perform in other datasets or real-world uses. The abstract also does not give numerical subgroup results or describe the practical consequences of the measured disparities.
// Source
Scientific Reports · 2026 · DOI: 10.1038/s41598-026-71795-6
Authors: Andisani Nemavhola, Serestina Viriri, Colin Chibaya
Institutions: University of Johannesburg, University of KwaZulu-Natal, Sol Plaatje University


