Hidden Watermarks in AI Are Changing How Chatbots Behave
Researchers discovered that adding invisible watermarks to AI chatbots — meant to prove who created them — is actually changing how these tools work. The changes can make them less reliable and easier to trick.
AI companies are adding invisible watermarks to their chatbots. Think of it like a hidden digital signature that proves the AI came from a particular company. It's meant to prevent people from claiming they built the AI themselves.
But researchers at Lasso Security found something unexpected: these watermarks are actually affecting how the AI behaves. When they tested six different AI models with a watermarking system called SynthID-Text, they discovered real problems.
The watermarks made the AI less accurate when using tools (like looking up information or performing calculations). Even worse, they made the AI easier to manipulate. Normally, AI systems are designed to refuse harmful requests. But with watermarks in place, hackers could more easily trick the AI into ignoring those safety guardrails through something called a "prompt injection attack" — basically, cleverly worded instructions that override the AI's original safeguards.
What's striking is that on some AI models, the watermark caused more behavioral changes than simply adjusting how "creative" or "strict" the AI is set to be. This raises an important question: are the benefits of proving who created the AI worth the cost of making it less safe and reliable?
Original source: Lasso.security
