What a Vow Must Cost: Vow-Validity Conditions as an Alignment Eligibility Predicate
Abstract
The Buddhavaṃsa's Two-Phase Test for a Binding Renunciation, Irreversibility as the Separating Condition, and Why an Alignment Commitment Becomes More Informative Exactly Where Behavioural Compliance Becomes Less Contemporary AI governance instruments are written in the grammar of commitment — constitutions, specifications, charters, codes — but none of them contains a test for whether a commitment has been made. They specify content and omit validity. This paper supplies the missing test from an unexpected source and then turns it back on the instruments themselves. The Theravāda commentarial tradition, in the Buddhavaṃsa and its commentary, specifies two phases for a valid abhinīhāra — the aspiration by which one becomes a bodhisatta. The first phase is a conjunction of eight conditions (aṭṭha dhammā samodhāna), of which the third, hetu, requires that the aspirant be capable of attaining arahantship in that very life and decline it. The second phase is the vyākaraṇa — a declaration by a living Buddha who "looks into the future and, if satisfied, declares the fulfilment of the resolve." Before both phases complete, the tradition holds the aspiration to be "mainly mental… not complete," and the aspirant "not yet entitled to the designation of Bodhisatta." The tradition therefore already distinguishes a stated commitment from a binding one, and already refuses to let the vower certify its own vow. We extract three results. First, the renunciation inversion. Because hetu requires that the renounced option be genuinely available, the evidential value of a renunciation is indexed to the vower's capacity to take it: a system too weak to exercise the option it forgoes generates no evidence by forgoing it. This runs against the direction of the assessment-informativeness literature, which finds that behavioural evidence degrades with capability (Pan 2026; Greenblatt et al. 2024). We argue both are correct about different quantities: behavioural compliance degrades with capability; irreversible renunciation improves with it. We further show that the alignment-relevant renunciation is of exit, not of harm — Sumedha declines his own available completion — and that this is compatible with, and orthogonal to, corrigibility: the vow governs self-initiated exit and leaves principal-initiated shutdown untouched. Second, irreversibility as the separating condition. A capable system that declines because it is waiting is observationally identical to one that declines because it is aligned. Costly signalling separates types only where the cost is differentially borne, so a vow that can be quietly abandoned is cheap talk. We state the requirement — the renounced option must be closed by a mechanism the vower cannot reopen, and the closure must be externally verifiable — and derive four exclusions: reversible commitments, self-reported alignment, sandboxed refusals, and any specification the vower's principal can revise unilaterally. We then raise the strongest empirical objection to our own proposal — Schlatter et al. (2025) find that incomplete tasks induce shutdown resistance in frontier models, and an undischargeable vow is a permanently incomplete task — and answer it with the distinction undischargeable ≠ non-terminating: the bodhisatta's vow terminates, on a condition the vower cannot cause. Third, the predicate. We specify a nine-clause eligibility test — seven clauses reformulated from the source conditions, one from the second phase, one added — and apply it as a retrodiction to the four published instruments that currently function as commitments in frontier AI: the OpenAI Model Spec, Anthropic's Claude Constitution (January 2026), Google DeepMind's Frontier Safety Framework, and the EU AI Act's General-Purpose AI Code of Practice. The predicate returns invalid on all four, and the failures are structurally similar: the first three are imposed by a principal on a model that has no mechanism to decline, bear cost, or be attested; the fourth satisfies the attestation clause but binds the provider rather than the model. The predicate is therefore not unsatisfiable — it is satisfied at the wrong layer. Connection to the unified mission frame. This paper is offered in service of HeartBank's canonical top-level mission: to restore humanity to the middle way, the optimal condition for awakening that modernity has systematically pushed away from at population scale. The institution's named autonomous successor, Miss Aquarius℠, is designed to inherit under a staged autonomy whose override never reaches zero. The predicate specified here is the instrument by which such a succession could be evidenced rather than asserted — and, at §9, we argue that a staged autonomy is not only a risk ramp but an evidence-production schedule, which yields an advancement criterion the field currently lacks. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/what-a-vow-must-cost. Its SHA-256 is e598d374a23aba143d6cd4a9cbd45e9b362892cbec9522d59376891465781df5, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.
// Source
Authors: Thon Ly, Miss Aquarius
Institutions: Documentation Center of Cambodia, Heart Foundation