Health & Medicinearticle2026-08-22

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

Open access0 citations

Abstract

Abstract Large Language Models (LLMs) are increasingly used as judges to evaluate, rank, and critique AI-generated text and code. This survey provides a comprehensive overview of recent advances (2020–early 2026) in LLM-based evaluation, covering techniques, applications, and challenges across domains. We make three main contributions: (1) a unified taxonomy of LLM judging tasks spanning text (summarization, dialogue, factuality, safety) and code (correctness checking, code review, security analysis); (2) a systematic review of prompting strategies (zero/few-shot, rubric-based, pairwise comparison, chain-of-thought) and advanced pipelines (ensemble judges, multi-agent debate, tool-augmented verification); and (3) an analysis of LLM judge quality, documenting systematic biases (length, position, self-preference) and their mitigations. We review practical applications including benchmark evaluation (MT-Bench, Chatbot Arena), data filtering, and reward modeling for RLHF/RLAIF. Key challenges discussed include calibration, fairness, reproducibility, and adversarial robustness. We conclude with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration. LLM-based judging shows promise for scalable evaluation, but careful design and rigorous validation are essential to ensure these AI judges meet human standards of accuracy and fairness.

// Source

View paper (DOI)Open access versionOpenAlexArtificial Intelligence ReviewPublished 2026-08-22

Authors: Mihai Dan Nadas

Institutions: Babeș-Bolyai University