AI watermarks: how do they work and what do they prove?

AI watermarks are being introduced in text as well as images and videos, to create transparency around the use of AI-generated content.

On 2 August 2026, Article 50 of the EU’s AI Act came into force. This requires generative AI companies, such as Anthropic, Google, Meta and OpenAI, to provide all text generated by their AI tools with a machine-readable AI watermark. This must be in place by December 2026. 

This isn’t law in the UK, but if your organisation has activities and outputs in the EU then it may still be affected. The UK Government is looking at making changes in the law too. 

Why do we need AI watermarks?

The innovation in AI watermarks has a dual purpose: to create transparency around the use of AI-generated content, and to increase trust in the tools and processes that create it.

In terms of academic integrity, watermarks promise even greater potential. But do they live up to that promise?

The first thing to say is that this new technology in watermarking text is at a very early stage. Take Anthropic’s new Claude watermarking update. The tool to detect the watermark has not been released yet, but Claude users are already developing and sharing workarounds online.

What is an AI watermark?

To identify AI-generated text, academics currently rely on detection tools such as Turnitin, which flag specific patterns and phrasing associated with generative AI use.

Because they rely on these subjective markers, these tools can raise false positives: for example, if the human author of a text is neurodivergent, or a second-language English speaker. In a research context, this can lead to false, discriminatory accusations against students and academics, while genuine offenders might escape detection.

AI watermarks are not intended to be a failsafe security measure: in fact, that’s impossible.

In contrast, watermarking is embedded in the text generation process. An AI model builds a sentence one word at a time. Each new word is chosen according to probability. In Claude’s case, there’s an extra element: the model generates a list of words with similar probability and chooses between them at random. This is intended to more closely mimic the human creative process.

Watermarking streamlines how AI models make these small choices, creating a consistent pattern that can be detected using the dedicated key. 

Understanding the limitations

AI watermarks are not intended to be a failsafe security measure: in fact, that’s impossible. Because the watermark is embedded in the structure of the text, significant editing or paraphrasing can break the pattern and make it undetectable.

Where the watermark is present, it only indicates that generative AI has been used, not how. A human-written text rephrased with ChatGPT may end up “watermarked” in the same way as a text generated using a prompt. 

This means that watermarking is a regulatory measure, not a comprehensive detection tool for academic dishonesty. However, the presence of an AI watermark should be a reliable indicator that generative AI was involved at some stage of the process.

ChatGPT open on a laptop

Watermarking in other media

We’re more used to seeing watermarking when it comes to images and videos – where a logo or text is embedded into the image. They’re often used to protect copyrighted images, so it’s important not to try and circumvent it by cropping, and any published work should be checked for watermarks in images and that those images have been correctly licenced, not just grabbed from Google Images. 

A lot of image and video editing tools export with a watermark on their free tiers, but more and more have AI integrations, so again AI involvement is possible in these cases.

Pixelshrink: transparent human communication

Changing tone or register, summarising dry facts or targeting a specific audience – this is where many people, including academics, are tempted to turn to AI models such as Claude. These tools can be very useful, but it’s increasingly worth logging where and how you’ve used it, so you can respond if someone asks how AI use affects your work’s integrity.

At Pixelshrink, our human authors communicate your expertise in authentic, compelling content that’s also SEO optimised, helping you reach the widest audience and maximise impact. Explore our content creation plans or get in touch today to discuss your needs. 

Share 'AI watermarks: how do they work and what do they prove?'

More from the blog

Beyond the press release: creative ways to announce research findings

Beyond the press release: creative ways to announce research findings

When your research uncovers something worth sharing, the next step is getting it out there. You’ve found the patterns. You’ve gathered the evidence. Now you need to get it in front of the people it can help… or the people who can help you take it further. What are the...

Subscribe to our Maximum Impact newsletter

Want to make an impact too?

Leave us your details and we’ll be in touch to have a chat about how we can help.

Contact Form - small (bottom of most pages)