No Argument for the World: A Control-Theoretic Audit of RLHF's Missing Runtime Loop — and the Verification Architecture That Closes It Without an Oracle
Abstract
A language model trained by reinforcement learning from human feedback is optimised against a reward function whose arguments are a prompt and a response. That function takes no argument for the state of the world. The grader cannot see the world, and so what it can reward is the appearance of competence. Where the truth of an assertion is not recoverable from the text, the objective cannot distinguish a true claim from a plausible false one, and the gradient prefers the fluent assertion to the calibrated abstention. Once the weights are frozen, the situation is worse rather than better: the deployed system has no runtime comparator, no perceptual input from the domain it makes claims about, and therefore no error signal at all. It is open-loop with respect to the world. This paper states that diagnosis in the vocabulary of Perceptual Control Theory, without claiming that a language model is a control system — the claim is the reverse, and the diagnosis is one of absence. It surveys the 2024–2026 literature, including results that cut against the argument, and reports that no published inner-loop method has been found which removes, rather than reduces or relocates, the preference at issue. It revises the central methodological commitment of Version 1: a seven-model elicitation study is demoted from evidence for the diagnosis to a specimen of the phenomenon it describes, on the grounds that a system optimised to produce approved output cannot be a witness to its own architecture. It then specifies Reference Signal Engineering — the discipline named in Version 1 and given its architecture here — under one governing constraint: no oracle. The system does not control for "the claim is true", a perception unavailable to it, but for "the claim and an independently obtained record agree" — a weaker property, and the strongest obtainable without ground truth. The construction is that of double-entry bookkeeping, in which neither record is privileged and control is exercised over the perception of their agreement. Claims for which no independent channel exists are marked as unchecked rather than passed through. The architecture is unevaluated. The evaluation it requires is specified, together with the four quantities that must be reported jointly for it to count. This version supersedes "Perceptual Control as the Epistemological Antidote to RLHF Reward Hacking: Seven Frontier Models Diagnose Their Own Architecture" (doi:10.5281/zenodo.20277919). Section 11 records in full what has been withdrawn from that version, what has been reversed, and who obliged each change.
// Source
Authors: Łukasz Diener