Since August 2026, new Claude models have been weaving an invisible AI watermark into their text. The wider industry is moving in the same direction on the transparency of AI-generated content, with around 190 organisations having signed the EU Code of Practice. For universities, this sounds like the answer many have been waiting for: a new signal that can help identify whether a piece of text came from an AI system.
The development is useful. It simply answers a different question from the one examination boards actually face.
Table of Contents
ToggleWhat an AI watermark actually does
Language models choose words one by one from several possible candidates. Whether “The weather today was cold and …” continues with “grey” or “overcast” is normally influenced by a random sampling process. An AI watermark changes that process in a controlled way. A secret key influences which statistically equivalent word choices are favoured. To the reader, the wording remains ordinary. To someone with the appropriate detection method, the resulting pattern can become statistically recognisable. Nothing is added to the visible text and there are no hidden characters. Technically, Anthropic describes its approach as related to SynthID Text, the watermarking method introduced by Google DeepMind in Nature in 2024.
Anthropic names the reason itself: the transparency obligations of the EU AI Act, which apply from 2 August 2026. That also tells us what the technology is meant to achieve. It is a mechanism for making AI-generated content more transparent and detectable. That is not the same thing as proving authorship in an examination.
The limits of Anthropic's watermark
Anthropic is unusually explicit about those limits. That list matters for any university considering whether an AI watermark should form part of an assessment or academic integrity process:
- It indicates that Claude was involved. It cannot establish whether Claude produced most of the text or was mainly used to revise it.
- It does not prove that a text was written by a human. It also cannot identify another AI system that uses a different watermarking method or no watermark at all.
- Short passages are more difficult to assess reliably because there are fewer word choices from which a statistical signal can emerge.
- Fact-heavy text and simple proofreading can leave too little variation for the watermark to provide a strong signal.
- Completely rewriting the text can remove the watermark.
Anthropic is particularly clear about what the watermark does not establish: ownership or authorship. That distinction matters. The question in an academic assessment is not simply “Was AI involved?” It is “What did this student actually contribute?” On that question, the watermark has little to say.
Why AI watermark circumvention is not a fringe issue
There is also a broader research problem. At ICML 2024, a team from ETH Zurich showed that watermarking schemes can be approximately reconstructed by querying a watermarked model. In the schemes they studied, this could be done for less than 50 dollars, with average success rates above 80 per cent for both removing and forging the watermark. There is an important qualification here: the study examined earlier watermarking schemes, not Anthropic's current Claude implementation. It does not demonstrate a direct attack on Claude's watermark. It does demonstrate that watermarking creates a technical arms race of its own.
The pattern is already familiar from AI detection. An entire market of “humanizer” tools has emerged with the specific purpose of changing AI-generated text so that detection systems are less likely to recognise it. There is little reason to expect watermarking to remain untouched by the same kind of adaptation, particularly when the provider itself describes complete rewriting as a way to remove the signal.
The deeper problem is authorship
Anyone who anchors academic integrity entirely in the finished product will eventually run into the same problem. A finished text can be changed. That has always been true of string matching in plagiarism detection. It is true of stylistic signals in AI classification. It is now true of statistical signatures in AI watermarking as well.
What is much harder to recreate afterwards is the path that produced the text: drafts, revisions, research activity, corrections, abandoned ideas and the development of an argument. That is why we consider writing process analysis a more durable form of authorship evidence. It does not depend on which model a student used, whether that model applies a watermark, or whether the final text has been rewritten afterwards.
None of this makes the AI watermark useless. Quite the opposite. As an additional signal for the identification and labelling of AI-generated content, it has real value. That is exactly how it should be used: as one indication among several, within a corroboration standard where no finding rests on a single tool. For transparency in the information ecosystem, it is a meaningful step forward. For the question of who is responsible for an academic work, it remains only one part of the evidence.
Making personal contribution visible
If you want to explore how authorship can be evidenced independently of the model used and of any AI watermark: see our toolset at a glance, or get in touch directly.






No comment yet, add your voice below!