AI & Computingpreprint2026-08-02

Cross-Model Behavioral Fingerprinting for Backdoor Detection in Enterprise AI Agent Supply Chains

Open access0 citations

Abstract

This work presents CloudGuard, a model-agnostic security framework for detecting backdoors and persistent state manipulation in enterprise AI agent supply chains. The framework extends behavioral fingerprinting beyond stateless tool execution by integrating software provenance, capability-normalized execution traces, data-flow and privilege transitions, conversational context, persistent memory, retrieval events, claim–evidence relationships, and downstream external effects into a Stateful Behavior-Provenance Graph. CloudGuard addresses heterogeneous enterprise environments in which AI agents may use different large language models, tool frameworks, Model Context Protocol servers, skills, plugins, retrieval systems, and long-term memory components. Its security controls include supply-chain provenance verification, write-time memory admission, read-time state re-evaluation, taint propagation, counterfactual influence analysis, open-set risk calibration, three-way ALLOW–STEP-UP–BLOCK enforcement, and signed behavioral attestation for runtime drift detection. The accompanying controlled evidence package contains two reproducible synthetic mechanism benchmarks. CMBF-Sim evaluates cross-model-style behavioral invariance under tool aliases, timing differences, retry behavior, telemetry loss, provenance manipulation, and previously unseen attack families. CMBF-State evaluates distributed conversational triggers, persistent memory poisoning, context-compaction poisoning, false-precedent insertion, procedure-memory manipulation, and trigger-conditioned semantic steering across linked write, persistence, retrieval, influence, and consequence stages. All reported experimental results are explicitly limited to controlled synthetic settings. They demonstrate internal mechanism behavior and failure boundaries, but do not establish production detection accuracy or effectiveness on real enterprise deployments. The manuscript therefore separates reproducible mechanism evidence from future validation requirements involving real multi-LLM executions, public agent-security benchmarks, independently governed enterprise data, adaptive red-team testing, and analyst-centered evaluations. Authors: Leo Lv(Lyu Chang) and Felix Fu(Fu Qiang)Affiliations: FLYINGNETS PTE. LTD., Singapore; Flyingnets株式会社, Tokyo, Japan.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-02

Authors: Chang Lyu, Qiang Fu

Institutions: FlyingBinary (United Kingdom)