Researchers randomly assigned physicians either to a control group or to a group that could use GPT-4o while working through standardized clinical vignettes. Reported score gains were largest in Kenya, followed by Indonesia and the Netherlands, and the differences in scores between Kenya and the Netherlands became smaller.

The results also varied within countries. Score variation increased in Indonesia and decreased in Kenya, and greater use of the language model was associated with larger gains. Some physicians without access still scored higher than some physicians with access, indicating that performance was not uniform among users.