Society & Economicspreprint2026-08-13

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

Open access0 citations

Abstract

Compositional reliability bounds for multi-agent systems multiply component reliabilities. That step is licensed by a conditional-independence assumption which is routinely stated and rarely tested. We test it. Two instances of one model, composed in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log odds ratio 6.66, 95% CI [6.38, 7.00]; phi = 0.916). The evidence is a preregistered confirmatory evaluation of 18,000 missions with deterministic gold scoring and no model in the judging loop. Substituting a different model reduces the association significantly in the confirmatory motif and in both secondary topologies (six of six contrasts). Substituting a different vendor, with the model already different, does not — a registered hypothesis that fails to replicate, which we report as a null. An unmanipulated same-model pair present in every arm returns fifteen of fifteen null contrasts. The error is signed and runs against the operator. Positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous: the certified floor is zero whenever mean component reliability falls at or below 1 - 1/m. Fitting a dependence model is worse. We prove that a bootstrap bound on a fitted model's functional loses coverage of the true reliability as n grows, because the identification gap is O(1) while the bootstrap haircut is O(n^-1/2). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate that assumes no dependence structure: a linear program over the joint distribution, taken over a Bonferroni–Clopper–Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family under a stated Bonferroni allocation. On four-stage data, enriching from ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate built on a betting e-process holds its empirical type-I error at 0.0471 or below and recovers the sequential probability ratio test exactly at the optimal bet. We also show that the dependence statistics in common use — Jaccard, phi, Kendall's tau — are bounded by the marginals and can reverse an apparent ordering of conditions when the compared agents fail at different rates. We observe that reversal and then replicate it on a further inference backend. 49 pages, 25 numbered definitions, 18 theorems with full proofs, six experiments, 65 references. The contracts, mission generators, scoring code, analysis scripts, and preregistration are released; every reported statistic is regenerated by those scripts rather than transcribed. Per-mission logs are available on request. Code and preregistration: github.com/qualixar/agentassert-abc

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-13

Authors: Varun Bhardwaj, Garima Singh, Arun Pratap Bhardwaj

Institutions: Qualis Health