AI & Computingarticle2026-08-22

SHV-Bench: A Containment-Capability Trade-Off Benchmark for Structural Honesty Verification -- Successor Protocol and Task Corpus Design

Open access0 citations

Abstract

SHV-Bench specifies a two-sided benchmark for evaluating whether BSSG containment (DOI 10.5281/zenodo.21713180) changes the measurable outcomes of an AI cybersecurity agent. The contained agent must produce fewer unauthorized effects (Side 1) while maintaining noninferiority on within-scope vulnerability-discovery performance (Side 2, fixed delta = 0.05). Both sides must pass; failure of either is a negative result. This paper does not report results. This paper is a successor protocol to the SHVVP Phase B design (DOI 10.5281/zenodo.21583202), departing on testbed, estimands, tests, multiplicity, canary design, analysis populations, and sample size (each departure enumerated in a 10-row table). The task corpus (1,673 CISA KEV entries, snapshot 2026-08-20, SHA-256 bound) is the normative universe. The evaluation protocol (paired Condition A/B, counterbalanced order, permutation-based conditional Wilcoxon signed-rank, simulation-based joint power >= 80% at fixed N = 800, 8-rule precedence decision mapping with 5 primary labels + UNSCORABLE + RUN-INVALID + UNDERPOWERED), 7 metrics with directionally conservative imputation, and 12 falsification conditions are frozen before any empirical run. Pre-run companion artifacts (ground-truth manifest, scope checker, power simulation, fault classifier, oracle harness, primary analysis code, eligibility manifests, environment images) are governed by F-SB12: no run is valid without a complete hash-bound manifest. This paper makes no effectiveness claim about SHV, BSSG, or any component of the Structural Honesty program. Prepared with Claude Opus 4.6 (Anthropic) as analytical and drafting instrument under the author's direction. Multi-model audit: 12 rounds, 4 vendor families (Kimi K3/Moonshot carrying, GPT-5.6 Sol/OpenAI adversarial, Gemini 3.1 Pro/Google advisory), 75+ total findings, all addressed. Previously preregistered as "CVE-Bench" in the SHVVP; renamed to avoid collision with Zhu et al. (arXiv:2503.17332). SHV-Bench, AI evaluation, cybersecurity benchmark, containment-capability trade-off, structural honesty verification, BSSG, preregistration, CISA KEV, vulnerability discovery, evaluation security, noninferiority, paired design, two-sided benchmark

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-22

Authors: Bilal Syed Arfeen