From lectures to learning outcomes: meaningful integration of AI-generated content in pre-clerkship medical training
Abstract
Abstract Background Large Language Models (LLMs) can efficiently synthesize educational content, yet few studies have evaluated standardized, LLM-powered curricular interventions and their effects on medical student learning. Methods In this single-institution study, lecture-specific summaries and Anki flashcards for the pre-clerkship curriculum were generated using LLMs optimized through a structured, content-agnostic prompt engineering process. Outputs were evaluated using standardized human grading for (1) hallucinations, (2) coverage of faculty-defined learning objectives, and (3) flashcard information density. The highest-performing configurations produced materials deployed across two 3-week blocks (genetics and pharmacology). Impact was assessed using block exam performance and a post-intervention survey in a resource-rich environment with existing faculty-created, commercial, and student-curated materials. Because AI use was self-selected, we compared baseline performance on the three preceding block exams and tested for a catch-up effect using difference-in-differences and analysis of covariance (ANCOVA). Results Optimized prompts yielded zero hallucinations in summaries and 1 per 21 flashcards (4.8%), with 100% mean coverage of faculty-defined learning objectives. Exam performance did not differ significantly between summary users and non-users in either the genetics (p = 0.76) or pharmacology (p = 0.35) block, and between-group differences fell within a prespecified equivalence margin in adequately powered comparisons. Flashcard use was likewise not associated with significant score differences (genetics p = 0.86; pharmacology p = 0.05), while AI-flashcard users had lower baseline performance than non-users (Cohen’s d = -0.36 to -0.59). Most users reported time savings with AI flashcards (74%) and summaries (61%) and rated them more time-saving than the platform’s default GPT summary (91%); greater use correlated with higher perceived utility for faculty notes and student flashcards (r² = 0.55, 0.79). Conclusions Pre-clerkship students who used AI-generated content showed no significant difference in exam outcomes, with differences within a prespecified equivalence margin in adequately powered comparisons, while reporting time savings and usefulness; baseline data suggest these users began with lower prior performance and caught up, though small subgroups (e.g., pharmacology flashcards, n=14), regression-to-the-mean effect, and near-universal genetics-summary exposure (92 of 98) limit inference. AI-generated content may offer an efficiency for programs producing materials and an opt-in study option within a well-resourced curriculum. Future work will examine less structured settings, such as clinical and surgical education.
// Source
Authors: Jay Khurana, Hossam A. Zaki, Ellie Pavlick, Jillian Turbitt, Heather McGee, Sahil Gupta, Sriya Sai Pushpa Datla, Salma Eldeeb, Thais Salazar Mather, Sarita Warrier, Joyce Ou
Institutions: Brown University