Human-in-the-loop Evaluation of medical large language models: A bilingual framework for clinical accuracy,saftey, hallucination detection , ARABIC-ENGLISH medical AI
Open access0 citations
Abstract
Large Language Models (LLMs) are increasingly being applied to healthcare-related tasks, including medical question answering, clinical documentation, medical translation, patient education, information retrieval, and healthcare workflow support. Despite their impressive linguistic capabilities, fluent model outputs cannot be assumed to be clinically accurate, safe, complete, or appropriate for a specific medical context.
// Source
View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18
Authors: HASAN AHMED RASHAM