The paper develops a taxonomy of watermarking methods based on when they are applied in the life cycle of a large language model, including before or during training and at the stages of token selection and sampling. It also connects the AI Act’s requirements with existing evaluations of watermark detectability, robustness and language-model quality.

Because interoperability has received little theoretical treatment in this research area, the authors propose three normative dimensions for assessing it. Comparing current methods with the four European criteria, they conclude that no approach yet satisfies all of them and recommend more research on watermarks embedded in the underlying architecture of language models.