NASA TLX workload profiles of multimodal artificial intelligence models and dental students during objective structured practical dental examinations
Abstract
Abstract Accuracy alone cannot characterize how artificial intelligence (AI) models engage with dental assessment tasks. This cross-sectional study introduces workload profiling as an evaluative lens, comparing modified National Aeronautics and Space Administration Task Load Index (NASA-TLX) dimensions elicited from ten multimodal AI models across three providers (Anthropic, OpenAI, Google) with those of twenty students on 70 objective structured practical dental examination items spanning four question formats (plain multiple-choice, visual short-answer, visual multiple-choice, and visual True/False). A two-step elicitation protocol separated task performance from workload self-rating; each model was tested in triplicate. Overall accuracy differed significantly across responders (χ² = 95.309, p < .001): students scored 77.9%, while models ranged from 44.3% to 74.3%. Intra-model reliability ranged from 0.677 to 0.955 (ICC, all p < .001). Seven models reported significantly higher mental demand than students, six reported higher temporal demand, and nine reported higher frustration. Convergence was greatest on plain multiple-choice questions and divergence widest on visual short-answer questions (χ² = 44.676, p < .001) where all models reported significantly elevated mental and temporal demand, yet only three showed higher frustration. Negative workload–accuracy correlations were notably observed in Claude Opus 4.6 and GPT-5.2 Pro. These findings indicate that AI models can detect task difficulty but show limited sensitivity to affective load, as reflected by attenuated frustration responses. NASA-TLX dimensions carried unequal validity when applied to AI, with workload ratings, particularly from Claude Opus 4.6 and GPT-5.2 Pro, showing promise as automated proxies for item difficulty estimation during examination development.
// Source
Authors: Sanaa N. Al‐Haj Ali, R. A. Farah