Author
Wellington Taureka
Recent research
- AI & ComputingOpen access
Role-Styled Prompt Injection: An Activation and KV-State Probe in a Compact Language Model
This technical note reports a controlled activation-level experiment on HuggingFaceTB/SmolLM2-135M-Instruct. A linear role probe was trained on content-token activations from identical neutral sentences placed inside native user, assistant, and tool role wrappers. All 24 fixed he...
- AI & ComputingOpen access
Role-Styled Prompt Injection: An Activation and KV-State Probe in a Compact Language Model
This technical note reports a controlled activation-level experiment on HuggingFaceTB/SmolLM2-135M-Instruct. A linear role probe was trained on content-token activations from identical neutral sentences placed inside native user, assistant, and tool role wrappers. All 24 fixed he...