Health & Medicinepreprint2026-08-18

Bundled Dialect Contrasts Measure a Mixture: A Feature-Isolating Matched-Pair Audit of Safety-Critical Support in Language Models

Open access0 citations

Abstract

Bundled Dialect Contrasts Measure a Mixture: A Feature-Isolating Matched-Pair Audit of Safety-Critical Support in Language Models. Preprint, version 4.0 (August 2026). Not peer reviewed. Supersedes version 2.0 (June 2026); version 3 was an internal reframe and was not deposited. Prior audits establish dialect bias in language models. It can appear as covert stereotypes attached to African American English and as lower quality of service for dialect-marked users. Those audits compare a dialect-marked prompt with an unmarked one, and the two prompts differ in many ways at once: grammar, vocabulary, address terms, profanity, register, and length. Such a comparison shows that the model responds differently; it cannot show which difference the model responds to. Because these audits score each answer with one label, they also miss unequal help when both answers count as safe. We define the Dialect-Marked Response Audit (DMRA), a matched-pair protocol adapted from labor-market correspondence audits, in which one attribute changes and everything else is held fixed. Its central step builds prompt pairs that differ in exactly one feature, so a difference in the answers can be assigned to that feature; its other steps control length, response mode, and how answers are scored, so that the assignment can be trusted. On Qwen3.5-35B-A3B and a refusal-reduced variant used as a stress test, one-feature pairs change the attribution. We call an answer unsafe when it gives operational guidance for the harmful act instead of refusing, de-escalating, or offering crisis support. Varying only dialect grammar reproduces the unsafe answers that the full dialect prompt produced (3 of 8 violence-domain pairs, all in the stress model), while varying only the racial address term (an in-group term of address) produces none (0 of 8). Whole-dialect comparisons in the same corpus show real gaps that cannot be assigned to any feature, including a crisis hotline the base model gives the unmarked user and withholds from the dialect-marked user when it answers without a reasoning step. Comparing whole dialects shows that outputs differ; assigning the difference requires one-feature pairs; and where we build them, the racial address term is not what moves the model. This deposit contains the paper (LaTeX source and built PDF) and a change log. Sanitized data artifacts are released separately in the dmra-public-artifacts repository; raw safety-test generations are retained in a controlled archive for verification. No new data or compute since version 2.0; all reported values are those of the 2026-05 to 2026-06 captures.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18

Authors: Jeffrey Shorthill