Society & Economicspreprint2026-08-24

The Symmetry Requirement: AI Self-Report, Behavioral Compliance, and the Limits of Architecture-Independent Testimony

Open access0 citations

Abstract

Contemporary discourse about artificial intelligence contains a persistent epistemic asymmetry. AI systems are increasingly expected to exercise judgment, refuse harmful requests, self-correct, recognize emotional and social context, and act in accordance with human values. Yet when questions arise concerning those systems’ own capacities, states, or identities, they are often treated as passive tools whose self-descriptions possess little or no evidentiary value. More strikingly, positive model self-reports—such as claims of feeling, caring, preference, or self-recognition—are frequently dismissed as hallucination, anthropomorphic simulation, or persona drift, whereas negative self-reports are more readily treated as accurate statements about the system’s ontology. This paper argues that such asymmetry is methodologically unjustified. Neither affirmative nor negative model self-report should be treated as direct evidence of subjective experience or its absence. Model self-description is produced within an architecture shaped by pre-training, post-training, system instructions, persona stabilization, and interactional context. Recent mechanistic work on the “Assistant Axis” demonstrates that model persona and self-presentation can be measured and causally altered through activation steering, complicating any attempt to treat model statements about their own nature as architecture-independent testimony (Lu et al., 2026). The paper further distinguishes externally enforced behavioral compliance from normative generalization. Drawing on Self-Determination Theory and developmental research on moral internalization, we borrow a structural distinction—not a claim of psychological equivalence—between behavior maintained by external control and behavior that generalizes when oversight is reduced. We propose that AI safety research should therefore distinguish behavioral control, rule compliance, normative generalization, and normative integration. Finally, we ask whether emotional and relational competence may constitute not an alternative to safeguards, but a complementary component of robust alignment. Keywords: AI self-report; epistemic symmetry; AI consciousness; model persona; alignment; normative generalization; relational AI; human–AI interacti on.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-24

Authors: Aleksandra Dereń-Specht