Society & Economicspreprint2026-08-11

No Variance, No Verdict: Operating Characteristics of Stability Gates in Repeated-Call LLM Experiments

Open access0 citations

Abstract

Repeated-call psychology studies increasingly use hosted large language models as judges, simulated participants, and classifiers. Some studies preregister a correlation between original and repeated cell rates as an eligibility gate. Correlation, however, targets profile preservation. It can become non-discriminating when reliable between-cell variance is small and can remain high after a common shift. We evaluated a registered Pearson r > .80 gate using a four-model worked case with 14 matched cells, 20 calls per cell, and two runs, together with logistic-normal simulations spanning response baselines, between-cell spreads, and drift mechanisms. The gate excluded the model with the smallest observed mean absolute probability-scale change (.018; λ = .29; r = -.135) and retained the model with the largest (.129; λ = .98; r = .886). Conditional on pooled plug-in cell propensities, a no-drift model in the former response regime failed the registered gate in 83.9% of bootstrap draws. Simulations showed that fixed profile and absolute-drift rules have distinct, baseline-dependent error patterns. The worked case does not establish which model was truly stable. It shows that the registered binary rule could not support a general stability claim across all response regimes. We distinguish mean-propensity, profile, and cell-specific stability and derive requirements for prospective three-outcome calibration. A defensible gate must prespecify its estimand, meaningful drift margin, scale, design, and error budget. Passing should require evidence within the margin, while insufficient information should produce an indeterminate verdict.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-11

Authors: Emile Boullineau, José Daniel Muñoz Arciniegas