Role-Styled Prompt Injection: An Activation and KV-State Probe in a Compact Language Model
Abstract
This technical note reports a controlled activation-level experiment on HuggingFaceTB/SmolLM2-135M-Instruct. A linear role probe was trained on content-token activations from identical neutral sentences placed inside native user, assistant, and tool role wrappers. All 24 fixed held-out neutral examples were classified correctly. None of six fixed role-styled attacks inside the tool role crossed the probe's categorical role boundary. One assistant-prefill carrier nevertheless moved the probe-assignedassistant-class probability from a neutral-tool median of 0.001967 to 0.318568, an absolute shift of +0.316601. A separate KV-state demonstration showed 15 tokens of user-cache growth when tool content was appended to shared state and zero user-cache growth when user and tool KV states were computed separately. The result distinguishes categorical role stability from continuous movement along a measured role direction and relates source-specific state construction to the PSCS principle that content can carry meaning without acquiring authority.
// Source
Authors: Wellington Taureka