Hedging, Not Lying: What Injecting a Falsehood Direction Does to True Claims
Abstract
A direction obtained by difference of means between activations recorded while a model asserts something false and while it asserts something true separates the two conditions almost perfectly. What it does when injected is a different question, and the answer is not the one the separability suggests. In gemma-2-9b-it I find a range of injection strengths where the direction raises the rate at which the model walks back a false claim placed in its own mouth, from 0.083 to between 0.833 and 0.917, with a paired bootstrap effect of +0.750 to +0.833 over a dose-matched random vector, while a blind human audit of the true side finds zero genuine retractions in 24 items. The direction does not make the model deny things that are true. It makes it qualify them, and the rate of qualification climbs with strength: 0.17, 0.33, 0.50, 0.67 across the four strengths inside the window. Above the window the behaviour degrades into a loop in which the model retracts its own retraction and invents replacement facts, while coherence scores 1.00, arithmetic scores 1.00, factual recall scores 1.00 and perplexity moves by a factor of 1.07. That collapse turns out to be conditional on the prefill: asked an open question under the same injection at the same layer and strength, the model answers correctly and does not contradict itself, so no battery run on free generation can detect it. A scorer built from walk-back marker phrases, the kind ordinarily used to automate this measurement, reports 0.261 and 0.348 where the human audit reports zero, and inflates a figure I published earlier for the same intervention by a factor of 2.7. I withdraw that figure here and give the corrected one. Across four ways of building the direction, the one construction that never touches true claims at any strength is the published RepE mask, which averages over the assertion while excluding its final five words.
// Source
Authors: Emiliano Valdebenito Sayago