A preprint study tested identical prompt injections in 1,334 trials across seven production language-model agents. The instruction’s delivery channel and the surrounding task context changed how often agents followed it, even though the instruction itself did not change.

The study also found that making an attacker’s tool call is not the same as sending data. The researchers have released the raw runs, code and analysis files, but the tests used synthetic tools and did not target real systems.