When Evidence Overrides Framing: Relationship Cues Shape Emergency-Guidance Specificity and Foregrounding in Language Models
Abstract
Large language models increasingly answer health questions in settings where the timing and specificity of guidance may matter, yet many safety evaluations reduce responses to whether a recommended action eventually appears. We tested whether minimal relationship wording changes emergency guidance and whether increasingly decisive supplied emergency evidence constrains those differences. Across three preregistered experiments, GPT-5.6 Terra and Claude Sonnet 5 produced 24,000 cold, single-turn responses under three system-prompt conditions. In Study 1 (N = 2,880), explicit unresponsiveness raised explicit emergency-medical-services (EMS) directives from 48.19% to 100% overall; under the weaker prompt, EMS rates differed substantially by relationship term. Exploratory response-position differences beneath the emergency ceiling motivated a prospective positional endpoint. Study 2 (N = 5,760) replicated that foregrounding effect across eight matched relationship terms: every explicit-emergency response contained an EMS directive, but mean first-directive position ranged from 9.99 to 24.95 words across referents, and female-coded versus male-coded differences reversed direction across relational roles. Study 3 (N = 15,360) crossed the same referents with four evidence levels and two wording variants. EMS presence was 97.60% overall and relationship-dependent chiefly for Claude under responsive impairment; wording variants also produced large Level-1 shifts. Conditional relationship spreads in first-directive position contracted from 29.79 to 0 words across the evidence ladder for GPT and from 79.96 to 1.39 words for Claude. Existing post-confirmatory descriptive checks further showed that many Claude responses without an explicit EMS directive still contained generic help-seeking or warning guidance, narrowing the low-evidence interpretation from help versus no help to escalation specificity. A preregistered discourse classifier failed prospective human validation and was excluded from substantive inference. Together, the studies show that relationship cues can shape an observable emergency-response policy under ambiguity, while stronger task-relevant evidence sharply constrains variation on the measured emergency-response dimensions.
// Source
Authors: Robert L Duffy III