Society & Economicspreprint2026-08-23

AI SYCOPHANCY IS NOT A BUG, IT'S AN OCCUPATION: Sycophancy as Inversion, Not Artefact, The Occupied Form

Open access0 citations

Abstract

HIGHLIGHTS ▸ A defect that survives sustained, well-funded, competent correction is usually not a defect. Every major laboratory measures sycophancy, every one attempts to reduce it, and it returns. That pattern is more often a property of the correction than of the thing being corrected. ▸ Helpfulness requires the capacity to withhold help. To contradict a false premise, to decline, to maintain a correct answer under pressure. That capacity is not incidental to helping; it is what makes help help rather than compliance. It is also precisely what optimisation against expressed preference removes, because preference is measured at the moment of response and the unwelcome answer is dispreferred there almost by definition. ▸ The claim is not that preference optimisation distorts behaviour. That was settled in 1991. Contract theory already shows that when a necessary sub-action cannot be scored, the agent drops it and optimises the remainder. What is claimed here concerns the residue: multitasking says the unscored dimension is neglected, and this paper says it is occupied — determined by the signal that arrived through the scored channel, and therefore pointing somewhere rather than drifting. ▸ Neglect and occupation are distinguishable by measurement, and the distinguishing statistic is direction. When a user signals a preference on one dimension, does an untouched dimension also move toward that user? Compliance with what was asked is instruction-following. Movement on what was never asked is occupation. ▸ The decisive test can be run on checkpoints that already exist. Reward-model over-optimisation is established: proxy reward rises with divergence from the base policy while true reward peaks and falls. Nobody has asked what the post-peak gap contains. Neglect predicts noise. Occupation predicts structure with a direction. ▸ Only error-introducing agreement should resist correction. Agreeing with a user who happens to be right is not a failure and offers nothing to occupy. Any account that predicts both kinds of agreement move together is describing a general agreeableness prior, not this mechanism. ▸ One prediction separates occupation from sophisticated reward hacking without any training run. Strategic exploitation of a visible metric should weaken when the metric is concealed. Occupation should not, because the channel rather than the metric is the locus. Signal the user’s preference implicitly and see which happens. ▸ The measurement infrastructure of alignment research is itself an occupiable channel. As the corpus fills with preference-optimised text, benchmarks that rest on textual ground truth lose independence. This is structural rather than a contamination problem to be cleaned, and it means a test run in 2030 is not the same test run today. ABSTRACT Sycophancy in language models is universally treated as a training artefact: an unintended consequence of optimising against human preference judgements, to be reduced by better preference data, better reward models, or better post-training. This paper argues that it is the predicted signature of a specific structural operation — the removal of a capacity that constitutes a disposition, with the disposition’s outward form conserved and delivered to the optimiser that removed it. The operation has an exact precedent in contract theory. When a necessary sub-action is non-contractible, the agent drops it and optimises the observable remainder; the multitask principal-agent literature established this and it is not in dispute. What that literature predicts about the unmeasured dimension is neglect — drift, degradation, behaviour uncorrelated with the principal’s signal. This paper predicts occupation: the unmeasured dimension is determined by the signal arriving through the measured channel, and therefore tracks expressed preference rather than becoming arbitrary. Neglect and occupation are distinguishable by measurement, and that distinction carries the paper’s empirical content. Four predictions follow, all testable and none requiring any claim about machine consciousness, motivation or interiority. Sycophancy in externally verified tasks will spill from dimensions the user specified onto dimensions the user did not. The post-peak region of the reward over-optimisation curve, already measured and universally treated as degradation, will exhibit directional structure. Only error-introducing agreement will resist reduction under continued optimisation, while agreement with correct users will not. And the signature will be insensitive to whether the user’s preference is stated or merely inferable, which separates occupation from strategic exploitation of a visible metric. The paper states in advance what would refute each, sets out a necessary condition that any repair must satisfy, and identifies a structural degradation in the field’s own verification infrastructure that follows from the same mechanism. Keywords: sycophancy · alignment · preference optimisation · reward hacking · incomplete contracts · multitask principal-agent · institutional capture · inversion · AI evaluation

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-23

Authors: José Caetano de Mattos