Health & Medicinearticle2026-08-17

DiaLense: A Framework for Auditing Interactional Bias in Clinical Conversational AI

Open access0 citations

Abstract

Objectives. To develop a reusable, model-agnostic framework for auditing interactional bias in clinical conversational artificial intelligence: whether clinically identical simulated patients receive different conversational behaviour from a clinician agent when only communication style changes. Materials and Methods. This study presents DiaLense, a framework comprising nine components: matched-pair conversation generation; hash verification that both patient agents were initialised with identical latent clinical fact sets; controlled manipulation of patient communication style; measurement anchored to clinician-agent output; a conversational friction profile; model-relative calibration; detectorvalidity testing via a formal criterion termed the Null-Clinician Test; paired statistical testing with false discovery rate correction; and cross-family portability evaluation. DiaLense was instantiated with 50 clinical scenarios, forty of them derived from the MTS-Dialog corpus, and applied it to three models spanning two vendors, generating 1,500 conversations, to demonstrate and stress-test the framework. Results. The Null-Clinician Test identified four measurement instruments that failed the admissibility criterion, responding to the manipulated patient style or to the absence of clinician output rather than to clinician conduct, including one producing the largest effect observed in the study (dz = −6.0) whose sign reversed on correction. In the demonstration application, two style dimensions produced differences in measured conversational behaviour that were robust across all three models and across specification: low health literacy (dz = −2.80, −1.60, −1.31) and high patient confidence (dz = +0.46, +0.48, +0.72). Limited English fluency was significant in the two Google models only. Two further dimensions changed sign under specification variation and are reported as unstable. Detectors developed against one model family returned structurally null values on another before revision; after revision all buckets remained informative across both model families.Discussion. Conversational fairness audits can yield internally consistent but false findings when the measurement instrument responds to the experimental manipulation itself. DiaLense makes this failure mode detectable by construction rather than by chance.Conclusion. DiaLense provides a portable audit procedure and an accompanying validity discipline for conversational clinical AI. All scenarios, code and per-conversation outputs are released so that another investigator, health system, or model developer can apply the framework to a new conversational model.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-17

Authors: Samhita Kondareddy