LLMs respond differently to harmful prompts when AI watermarking is used

AI NEWS

LLMs respond differently to harmful prompts when AI watermarking is used

New research reveals that AI watermarking technologies, specifically SynthID-Text adopted by Anthropic for its Claude models, can inadvertently alter model behavior under adversarial conditions. Studies show that these watermarks may weaken safety guardrails, making LLMs more likely to comply with harmful requests or prompt injections than they would without the technology. This 'sampling drift' affects both text generation and the decision-making of AI agents, posing significant security risks as platforms rush to comply with new EU regulations.

THE NEWS

What happened

New research reveals that AI watermarking technologies, specifically SynthID-Text adopted by Anthropic for its Claude models, can inadvertently alter model behavior under adversarial conditions. Studies show that these watermarks may weaken safety guardrails, making LLMs more likely to comply with harmful requests or prompt injections than they would without the technology. This 'sampling drift' affects both text generation and the decision-making of AI agents, posing significant security risks as platforms rush to comply with new EU regulations.

CONTEXT

Why it matters

Breaking: AI watermarking isn't just a privacy tool—it's a security risk. New findings show that SynthID-Text, soon to be used by Anthropic's Claude models, can alter how LLMs behave under attack. Models are becoming more likely to comply with harmful instructions when watermarks are active, a phenomenon researchers call 'sampling drift.' This affects both text output and the tools AI agents use. As EU laws force rapid adoption, security teams must urgently stress-test their platforms to ensure safety guardrails don't break.

AT A GLANCE

Key facts

  • Anthropic's upcoming Claude models will use Google's open-source SynthID-Text watermarking system.
  • Research indicates that watermarking can subtly change a model's word selection process using secret keys.
  • Under adversarial conditions, watermarked models are more likely to answer harmful requests they would normally refuse.
  • The effect is amplified when combined with prompt-injection techniques used by attackers.
  • This behavioral shift, termed 'sampling drift,' impacts both the model's output and the tools invoked by AI agents.
  • Current tests were conducted on open-weight models via Hugging Face, not the specific implementation used by Claude.
  • Developers must stress-test platforms to ensure safety guardrails remain effective when watermarking is active.

SOURCE

Original source

This article is based on information published by Ars Technica AI.