Separating Behavioral, Latent, and Causal Commitment to Incorrect Answers in Large Language Model Reasoning: A Structured Critical Meta-Synthesis
Abstract
The literature on multi-step chain-of-thought (CoT) reasoning increasingly claims that models become committed to incorrect answers mid-trace. We show that this shared vocabulary masks three non-interchangeable constructs—behavioral preference concentration, latent detectability, and causal irreversibility—and that evidence for one is routinely treated as evidence for another. Through a structured critical meta-synthesis over a core corpus of 23 studies (PRISMA 2020 reporting discipline; explicit search strings, eligibility criteria, and a thirteen-field extraction schema), we organize the evidence into behavioral, latent, and causal tiers and audit what each can license. The corpus supports a detection–manifestation gap but does not establish irreversibility. Quantified reporting gaps are themselves a primary finding: 0/6 probing studies report control-task selectivity; 1/16 intervention studies report post-intervention output coherence; and the only commitment boundary characterized as a single-step transition is estimated under greedy decoding with an answer-forcing suffix. Our contribution is a measurement framework: definitions for a point-of-no-return (PoNR_τ,w), a stochastic recovery curve R(k), normalized latent lead time (NLT), and behavioral preference concentration (BPC), plus a proposed intervention-competence control. Literature screened through 30 June 2026. Protocol, extraction matrix with per-cell provenance, and count-reproduction script accompany this paper.
// Source
Authors: Safwat Shabib, Prerona Arman