Health & Medicinearticle2026-08-18

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

Open access0 citations

Abstract

<title>Abstract</title> Background. Studies of artificial intelligence tools used in patient triage have largely involved academic medical center cohorts, scripted patient-actor scenarios, or knowledge benchmarks. Populations that may rely on such tools due to constrained access to in-person care, including Medicaid patients, have been less fully evaluated. Objective. To compare combinations of safety guardrails added to artificial intelligence tools for triage of patient-initiated text messages in a multi-state Medicaid population. Methods. Retrospective evaluation of 2000 messages from Medicaid patients across three U.S. states (Virginia, Washington, Ohio) during January 2023 through November 2025. Three physicians independently adjudicated each message under blinded review; disagreements were resolved by majority and a senior-physician arbiter. Tools included a deployed decision support system with a Conservative Q-Learning controller, supervised baselines (XGBoost+sentence-BERT, logistic regression, rule-based guardrails), and frontier large language models (Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, each with and without retrieval-augmented generation). Combinations spanned single tools, ensembles, cascades, multi-model consensus rules, and four-component guardrail compositions identified by a structured literature review. Thresholds and Platt calibration were fit on a held-out validation split and frozen before the test pass. Two pre-specified targets: an autonomous (physician-unassisted) triage benchmark of sensitivity and specificity both 0.80 or higher; and a sensitivity-floor target of 0.80 to 0.95 with clinician review of flagged messages. Results. Real-world messages had lower reading level (grade 4.590 versus 5.810) and more colloquialisms (59.1% versus 19.5%) than physician-scripted scenarios. Hazard prevalence on blinded physician review was 8.2% (165 of 2000). Two configurations met the sensitivity-floor target: a high-recall first-stage screen (sensitivity 0.855; 12.0 missed hazards and 729 alerts per 1,000 messages) and a disagreement-stratified clinician-review workflow (sensitivity 0.939; 5.0 missed hazards and 878 alerts per 1,000 messages). No configuration met the autonomous benchmark; sensitivity and specificity reached 0.594 and 0.592 for the best balanced single tool, 0.685 and 0.575 for the best balanced ensemble, and 0.297 and 0.874 for the highest-specificity cascade. Conclusions. No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting. Two configurations met the pre-specified sensitivity-floor target under clinician review of flagged messages, following the classical clinical-screening pattern.

// Source

View paper (DOI)Open access versionOpenAlexBMC Medical Informatics and Decision MakingPublished 2026-08-18

Authors: Sanjay Basu, Sadiq Patel, Parth Sheth, Bernardo Arevalo, Jeremy Schifberg, John Morgan, Rajaie Batniji

Institutions: University of California, San Francisco, Virginia Commonwealth University, California University of Pennsylvania, Highmark Blue Cross Blue Shield