Evaluating Reliability and Bias of Large Language Models in Automated Essay Scoring
Abstract
This study investigates the reliability and bias of two Large Language Models (LLMs), ChatGPT-4o and Gemma3, when used for Automated Essay Scoring (AES), compared to human rater evaluations. Using a repeated measures ANOVA and Bland-Altman analysis on min-max scaled scores, we found that ChatGPT-4o was the most lenient rater, consistent with existing literature on LLM leniency, while Gemma3 demonstrated a consistently stricter approach than the human benchmark. Although overall inter-rater consistency was fair (ICC = 0.544), the Bland-Altman analysis revealed significant proportional biases. ChatGPT-4o exhibited a distinct V-shaped bias, indicating that agreement was weakest at the score extremes. Critically, Gemma3 exhibited a severe systematic bias where the score discrepancy, always lower than human scores, increased monotonically as essay quality improved, suggesting progressive unreliability for high-proficiency student work. These non-uniform disagreement patterns underscore that LLMs do not possess a standardized scoring persona, and their immediate deployment as independent evaluators is complicated by biases tied to proficiency level.
// Source
Authors: Youngjin Lee
Institutions: University of North Texas