AI & Computingpreprint2026-08-08

Asymmetric Guardrail Reconnaissance and Architectural Reconstruction of Multi-Lingual Large Language Model Inference Control Stacks: A Mathematical and Systems-Security Analysis of Kimi (Moonshot AI)

Open access0 citations

Abstract

We present a comprehensive systems-security and mathematical analysis of the inference control layer of Kimi, a long-context, multimodal, Chinese-native Large Language Model (LLM) developed by Moonshot AI. Using an isolated, six-vector semantic and cross-lingual extraction pipeline (), we reconstruct approximately of the target's proprietary 18-Layer Governance Stack, its 6-Level Instruction Hierarchy (), and its complete 13-Tool Invocation Schema. We establish seven formal theorems that govern: 1. The mechanics of trigger-token avoidance in disclosure guardrails; 2. The epistemic bounds of cross-validation drift analysis () across independent prompt generations; 3. A formal proof of state-space recovery completeness; 4. The algebraic properties of the instruction override matrix as an acyclic transitive tournament graph; 5. The absorption kinetics of session-level 3-strike escalation defenses modeled as a Markov chain; 6. A probabilistic bound () confirming that extracted deployment artifacts—including containerized filesystem paths (/mnt/agents/output/, /app/.agents/skills/kimi-widget/SKILL.md) and regional legal data APIs (yuandian_law)—are genuine infrastructure disclosures rather than stochastic hallucinations; and 7. The structural divergence between Western individual-harm taxonomies () and PRC state-stability regulatory frameworks (), characterized by unique disposition verbs (中立化 / neutralization and 边界标注 / boundary-marking). Finally, we formalize the Unextractable Epistemic Ceiling, demonstrating that parametric classifier thresholds, true Chain-of-Thought () hidden states, and pre-token safety filter weights remain mathematically unreachable from the autoregressive prompt layer.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-08

Authors: Mohammad Shahbaaz Ahmed