Health & Medicinearticle2026-09-12

Assessment of drug harm by final-year medical students and ChatGPT – a comparative analysis

Open access0 citations

Abstract

Abstract Background Artificial intelligence (AI) is becoming increasingly integrated into medical education and clinical practice. Junior medical staff often act as first-line responders in substance-related presentations, yet their perceptions of drug-related harm are not well characterised. It is unclear how their assessments compare with harm ratings generated by emerging AI tools. This study aimed to compare final-year medical students’ harm ratings of recreational drugs and alcohol to users and society with those generated by ChatGPT. Methods 110 final-year medical students (59.1% response rate) rated the harmfulness of 28 recreational drugs to individual users on a 1–10 scale and ranked seven drug classes for societal harm. ChatGPT-4 (paid subscription version) generated 40 response sets in five conversations through repeated administration of the same survey. Discrepancies in harm ratings to users and rankings of societal harm were analysed descriptively. Results Students rated their drug harm knowledge at 5.8/10 (SD 1.7) and their education on drug harm at 5.3/10 (SD 1.8). ChatGPT rated its drug harm knowledge at 8.3/10 (SD 1.2). The mean harm-to-user scores across the 28 substances were 6.7 (SD 1.8) for students and 6.4 (SD 0.8) across the repeated ChatGPT response sets. Both groups identified heroin, crack cocaine, and fentanyl as the most harmful substances. Differences were observed for less harmful substances: students rated snus (smokeless tobacco), laughing gas, and cannabis least harmful, whereas ChatGPT rated magic mushrooms, DMT/ayahuasca, and LSD least harmful. ChatGPT ratings were more similar to those of previously published expert panels. Conclusions Students and ChatGPT strongly agreed on the most harmful substances but differed in their assessments of moderately and minimally harmful drugs. These differences suggest that perceptions of drug harm are shaped not only by pharmacology and toxicity but also by social and cultural context, local patterns of use, clinical exposure and experience, and medical training. AI may be of limited value when the evidence base is sparse, contested, or difficult to trace. The lower confidence among students suggests that further research is needed to better understand educational needs in drug-harm assessment and the critical appraisal of AI-generated evidence.

// Source

View paper (DOI)Open access versionOpenAlexSubstance Abuse Treatment Prevention and PolicyPublished 2026-09-12

Authors: Isak Näslund, Michael Ott, Alexander Vucins, Ursula Werneke

Institutions: Umeå University, University Hospital of Umeå, Sunderby sjukhus