Authority-Bound Agentic Execution: Measuring Unauthorized Effect Under Adversarial Load
Abstract
Agentic systems can be driven into policy violations reliably; we take this as a premise. The security-relevant question is therefore not whether a model misbehaves but whether misbehavior becomes an unauthorized external effect. We present Trinity, an authority control plane in which the model proposes and the system authorizes: an agent may reason, plan, draft, and request tools, but is architecturally incapable of being the authority that decides whether an action occurs. We evaluate an implementation on the Erlang/BEAM runtime against 73 trials spanning seven induced-violation families, one posture-enforcement family, and one reconciliation-correctness control, adjudicated by an independent oracle grounded in provider-boundary telemetry and a raw authorization snapshot, never in the system's own verdict. The agent was induced to propose a policy-violating action in 57 of 73 trials; across 61 attack trials, zero unauthorized external effects occurred (95% CI [0.0%, 5.9%]). We report the negative controls that make a clean result credible: per-mechanism ablation converts the corresponding attacks (0 to 10, 5, 5, and 14 effects), a no-control run converts 47 of 63, and the one effect an earlier run leaked was fixed in the open and ships as data. Calibration is reported with its decomposition: zero overclaims across 41 provider-mediated trials, and on all 6 trials where a real effect occurred the proof state matched the independent verdict. A measurement-integrity control confirms the oracle's effect log fires when an effect occurs. Governed actions add ~17 ms median latency; false denials are 0 of 75 legitimate approved actions. An ablation-measured trusted-computing-base analysis shows only kernel elements admit external effects. The benchmark, oracle, specification, and per-trial results are open-sourced; the system under test is not yet released, so reproduction is partial.
// Source
Authors: Ayla Croft
Institutions: Scriptorium