Health & Medicinearticle2026-08-22

Comparing the readability of AI-generated and society-authored patient information leaflets in orthopaedics

Open access0 citations

Abstract

Abstract Introduction Effective patient education is critical in orthopaedic care, influencing satisfaction, adherence, and outcomes. Artificial intelligence (AI), particularly large language models (LLMs), could offer the potential to improve patient information leaflets (PILs), but this remains underexplored. This study aimed to evaluate the readability of AI-generated orthopaedic PILs compared to UK professional orthopaedic society materials using objective metrics. Methods A retrospective quantitative study was conducted comparing PILs from nine UK orthopaedic subspecialty societies with matched AI-generated counterparts created using ChatGPT 4.5. AI responses were generated using simple, single lined patient-style prompts to simulate real-world queries. PILs were categorised as either condition-based, procedure-based, or general information leaflets. Readability was assessed using validated metrics including Flesch-Kincaid Grade Level (FKGL) and Reading Age, FORCAST, New Dale-Chall, SMOG, Gunning Fog Index, and Flesch Reading Ease (FRE). Word counts were also analysed. Grade levels were interpreted according to U.S. educational standards. Statistical comparisons between AI and human-generated materials were performed using appropriate parametric and non-parametric tests, with statistical significance set at p < 0.05. Results Across 134 orthopaedic PILs, AI-generated materials were consistently shorter in word count across all categories ( p < 0.01). Despite no significant differences in FKGL for General and Procedure based PILs, AI-generated Condition PILs demonstrated significantly higher FKGL (10 [9.2–10.6] vs. 8.7 [7.8–9.6], p < 0.01). FRE was consistently lower in AI-generated texts across all categories ( p < 0.01), suggesting reduced accessibility. AI materials also demonstrated significantly higher FORCAST and New Dale-Chall Grade Levels across all categories (all p < 0.01), indicating greater reading complexity. Conclusion AI-generated PILs offer brevity but do not consistently improve readability, with some indices suggesting increased complexity. While AI holds promise, clinician oversight and further validation are essential to ensure AI-generated materials enhance, rather than hinder, patient understanding and engagement.

// Source

View paper (DOI)Open access versionOpenAlexJournal of Orthopaedic Surgery and ResearchPublished 2026-08-22

Authors: Nasir Kharma, Saran Singh Gill, Chayan Shanmugaratnam, CM Gupte

Institutions: Imperial College London, Imperial College Healthcare NHS Trust