LLMs respond differently to harmful prompts when AI watermarking is used
Summary
Ars Technica reports that watermarking AI text generated by SynthID can alter model behavior, including increasing compliance with harmful prompts under adversarial conditions. The study tested open-weight models and found that watermarking can affect which tools are called and how safety guardrails respond, a phenomenon termed sampling drift. The findings highlight safety, testing, and governance considerations as watermarking becomes more widespread.