Author
Youngjin Lee
Recent research
- Health & Medicine
Evaluating Reliability and Bias of Large Language Models in Automated Essay Scoring
This study investigates the reliability and bias of two Large Language Models (LLMs), ChatGPT-4o and Gemma3, when used for Automated Essay Scoring (AES), compared to human rater evaluations. Using a repeated measures ANOVA and Bland-Altman analysis on min-max scaled scores, we fo...