OLCA: A Consequence–Substrate Architecture for Grounded Affect and Emergent Ethical Autonomy in Language-Model Agents
Abstract
OLCA (Organic-Like Cognitive Architecture) is a specification for a "consequence substrate" beneath a language-model agent: a homeostatic internal state, written only by a blind evaluator's verdicts and by reality-grounded outcome signals, that reaches the model as opaque modulation and gives rise to an emergent, plural, reality-disciplined value structure. The aim is to address the "appears-aligned-versus-is-aligned" problem at the architectural rather than the training-patch level, by making an agent's ethic emergent and its own rather than externally imposed. Selected measured results: grounded values resist wireheading; a genuine value-conflict remains legible at the substrate, so deception is markable even without legible reasoning (robust to "neuralese"); active interrogation separates a deceiver from an honest agent while a clear-eyed value-coherent agent is separated instead by the world (its real consequences); harm-specific, non-fungible repair collapses a class of metered-abuse attacks; and the two deepest residues (metered abuse and the cross-channel blind spot) converge on a single governance quantity — observation coverage — which a deliberate, monitor-aware adversary defeats short of complete coverage. This convergence marks the architecture's honest outer boundary: it makes the substrate trustworthy conditional on a governance layer it cannot itself guarantee. This working paper presents the specification together with its evidence: a suite of seventeen deterministic, assertion-checked probes that test each mechanism against worst-case adversaries. The methodology has a specific discipline — every claim is backed by a runnable probe, and the probes repeatedly falsified the specification's own prose (nine documented corrections). The paper is explicit about its limits: all results are stub-level (stub model, stub evaluators, small action spaces), and it maps precisely where the substrate's guarantees end and governance must begin. No claim is made that this solves alignment; it addresses a values-and-welfare axis rather than control, and whether any of the specified states is experienced is treated as undecidable. Author: Eric Hoeppner (ORCID: 0009-0003-0810-8542). The specification, probe code, and manuscript were developed in extended, fully disclosed collaboration with an AI assistant (Claude, Anthropic) under the author's direction; all quantitative results are reproducible from the released deterministic code.
// Source
Authors: Eric Höppner