Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study
Abstract
Multimodal large language models are increasingly used for dermatologic queries, but whether Fitzpatrick skin type (FST) labels affect diagnostic accuracy is unknown. We evaluated 656 biopsy-confirmed photographs from the Diverse Dermatology Images (DDI) dataset under no-FST, DDI-concordant FST, and two DDI-discordant FST conditions using ChatGPT 5.2 Edu and Gemini 3.1 Pro browser configurations (5248 evaluations). Confirmed malignant diagnosis omission from the top three differential diagnoses (“lethal miss”) was a prespecified exploratory outcome. Gemini had higher ordinal accuracy than ChatGPT (odds ratio 2.32; 95% confidence interval 2.01–2.67; false discovery rate-adjusted p < 0.001). No FST prompt condition significantly improved accuracy over no-FST prompting; an equal-image-weighted sensitivity analysis yielded similar results. Among malignant images eligible for extreme discordance, extreme DDI-discordant prompting was associated with more ChatGPT lethal misses than no-FST prompting (paired odds ratio 5.0; p = 0.043); Gemini showed no significant paired shift (p = 1.00). Post hoc binary malignancy detection showed low no-FST specificity for ChatGPT (42.5%) and Gemini (22.9%), indicating frequent overcalling. Gemini showed lower no-FST accuracy for malignant FST V–VI images than for FST I–II and III–IV images. FST prompting provided no measurable diagnostic benefit. Neither configuration supports autonomous dermatologic use.
// Source
Authors: Manoj Bhagwat, Tyler Wittles, Jeffery Tan, Joshua Mijares, Neil K. Jairath, Syril Keena T. Que
Institutions: Indiana University School of Medicine, Indiana University – Purdue University Indianapolis, New York University