AI & Computingarticle2026-08-18

Human-in-the-loop Evaluation of medical large language models: A bilingual framework for clinical accuracy,saftey, hallucination detection , ARABIC-ENGLISH medical AI

Open access0 citations

Abstract

Large Language Models (LLMs) are increasingly being applied to healthcare-related tasks, including medical question answering, clinical documentation, medical translation, patient education, information retrieval, and healthcare workflow support. Despite their impressive linguistic capabilities, fluent model outputs cannot be assumed to be clinically accurate, safe, complete, or appropriate for a specific medical context.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18

Authors: HASAN AHMED RASHAM