Society & Economicsarticle2026-08-22

Assessing AI-Mediated Interviewing Quality: A Theory-Grounded Comparison of Models for Evaluation Data Collection

0 citations

Abstract

As large language models are increasingly used in evaluative settings for their conversational potential in interviews, systematic frameworks for assessing their methodological quality are still lacking. Drawing on evaluation, qualitative interviewing, and validity theories, this study develops a theory-grounded framework for evaluating the quality of AI-mediated interviewing in evaluation contexts. This framework is then operationalized into eight assessment measures and applied in a convergent mixed-methods design to compare six generative AI models (five prompted-only and one fine-tuned) across two realistic evaluation interview scenarios. Findings reveal that AI models reliably detected and probed incomplete or irrelevant responses, and the fine-tuned model matched benchmark performance on nearly 70% of quality measures. However, performance was context-dependent, and models struggled with neutrality and clarification probing. The study contributes a replicable assessment framework, an empirical benchmark for AI interviewing performance, and guidance on appropriate deployment conditions and essential human oversight mechanisms.

// Source

View paper (DOI)OpenAlexAmerican Journal of EvaluationPublished 2026-08-22

Authors: Ali Safarnejad, Hippolyte Lefebvre

Institutions: University College Dublin, United Nations Children's Fund India