AI & Computingpreprint2026-08-22

Optimal, Not Perfect: A Test-Bed Architecture for Rule-Governed, Automatically-Auditable AI Advisory Systems, with Counselling as the Leading Case

Open access0 citations

Abstract

Can two separate evidence streams be composed into a rule-governed AI advisory system — with counselling as the leading case — that is, on specifiable axes, more consistent and more verifiable than either stream alone could specify? The question is narrower than whether a machine can be a perfect psychologist, a standard neither machines nor humans meet. Classical psychology and the clinical specialisation of Friction Theory say what intervention-mechanisms work on the human substrate; a younger literature on language-model substrates says how to make a model follow fixed rules, encode them robustly without rigidity, let present context override prior disposition, and — critically — where the model fails at trust-modulated responsiveness. The paper composes five such components into one architecture — not because silicon and the human client are the same kind of system, but because together they solve one operator problem: reliable, auditable, rule-governed advice. Its organising principle is optimal, not perfect, compliance: rule-following is itself a route-race, and both too-little and too-much commit-pressure on the rule-route are failure modes (unsafe or sycophantic advice on one side, rigid reactance-inducing refusal on the other). Because the product is wholly digital, this optimum is auditable rather than postulated: every rule and session can be audited automatically and at scale — behaviourally (was the rule applied) and for groundedness (is each answer supported by what was retrieved) — a coverage no human-delivered counselling can match at that cost, exhaustive for deterministic rules and sampled-human for the conditional ones. A deflationary safety working-hypothesis follows from the project's own pilot measurements: the model's own signals are an unreliable self-check, so reliable safety is placed in external coverage/grounding/confidence gates, a separate stronger verifier, and per-substrate and per-language re-calibration, rather than in the model's tuning. The contribution is positioned against the production retrieval-augmented-generation-plus-guardrails-plus-deferral stack, not only against therapy-architecture papers: a derivation whose design choices are motivated by a competing-directive mechanism and a rule-type split; an auditability spine that runs in deployment, per turn and exhaustively; and a gate-borne safety hypothesis. It is offered as hypothesis and invitation — an architecture and a test-bed, not a clinical product and not a claim of efficacy; the public, real-user demonstration is carried by non-clinical instances (a policy/handbook assistant and a tutor), while a clinical instance is treated as a regulated medical device and run only under research governance.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-22

Authors: Tomas Pødenphant Lund

Institutions: Aarhus University