Implementation of clinical grading committees in internal medicine sub-internships: implications for assessment accuracy and standardization
Abstract
The Internal Medicine clerkship has faced challenges in ensuring fair and accurate grades, with past studies examining areas including unconscious bias, inter-rater reliability, and variability. To address these issues, the Alliance for Academic Internal Medicine recommends the use of Clinical Grading Committees (CGCs) as a best practice to enhance grading equity and validity. However the use of CGCs in the Internal Medicine Sub-internship, a rotation which relies largely on subjective evaluation of clinical performance, remains underexplored. In this study, we examine whether CGCs reduce grade inflation, minimize site-specific variability, and promote grading consistency across multiple clinical sites of the Internal Medicine sub-internship rotation. We conducted a retrospective analysis of IM sub-internship grades for three academic years (2021–2024) at Hackensack Meridian School of Medicine across three clinical sites. IM students ( n = 103) were evaluated by faculty using a Clinical Evaluation Tool that rated 14 Entrustable Professional Activities on a four-point scale. Grades were calculated before CGC review (Pre-CGC) and after committee discussion (Post-CGC), utilizing a tiered grading system. A Chi-square test for independence was performed to assess grade inflation and inter-site variability pre- and post-CGC. There was no statistically significant difference between the mean calculated grades pre-CGC and post-CGC (Pre-CGC Mean = 3.37, SD = 0.75; Post-CGC Mean = 3.41, SD = 0.60; t (102) = -0.60, p = .55). Grade distributions across the three clinical sites also showed no significant variability pre-CGC (𝝌² (4, n = 103) = 3.55, p = .47.) or post-CGC (𝝌² (4, n = 103) = 0.73, p = .95). While our study did not support our initial hypothesis of CGCs reducing grade inflation, the lack of significant grade variability across sites limited our ability to accurately assess their role in standardizing grades. However the observed trends of post-CGC grades more closely approaching a mean may suggest that CGCs promote consistent grading practices across sites, likely due to integration of both quantitative and qualitative assessments, and the encouragement of more deliberative decision making. Although the study was underpowered, it contributes to an important and ongoing effort to improve fairness and accuracy of grading and assessment in medical education. Future work should include multi-institutional studies to further examine CGC impact on grading equity, perceived fairness, and learner outcomes in the sub-internship setting.
// Source
Authors: Deepti Reddy, Marcella Katsnelson, Chosang Tendhar
Institutions: UC San Diego Health System, University of San Diego, Center for Discovery