Petteri Pyyny
13 Aug 2026 3:17
Both the European Union and the state of California now require AI-generated content to be identifiable and clearly labeled as being created by AI.
So this is not merely an "annoying European Union requirement". An essentially identical requirement is also coming into effect in California, the most populous state in the United States.
And AI companies have already said that they are taking steps to comply with these requirements. Images, videos and audio are being given so-called digital watermarks that can be detected programmatically from the file, even though they are not actually visible in the image or video - or audible in the audio.
But what has caused considerably more confusion is the fact that AI companies are also adding digital identifiers to AI-generated text. For example, Anthropic has already stated that all text generated by its Claude AI will contain a digital identifier within the European Union starting August 2, 2026.
Many smug commenters may scoff at this and say that AI-generated text is easy to recognize anyway. But that is not actually the case. Text produced using clever prompts and more advanced AI models is, in practice, virtually impossible to reliably identify as AI-generated.
Yet both the EU and California require even text to be identifiable as having been generated by AI.
So how on earth does that work? Anthropic says in its own documentation that the identifier embedded in the text survives even if the text is copied to the clipboard and pasted as plain text into a text editor. According to the company, the identifier also survives if the text is edited (to a limited extent).
There has been plenty of discussion about the technology online, but the most likely explanation lies in a combination of cryptographic techniques and the underlying technology of AI models.
When an AI generates text, it selects successive words according to certain logic and probabilities. By slightly adjusting those probabilities according to a specific formula known only to the company, the relationships between individual words form a statistical pattern that deviates from the baseline probabilities. That pattern can then be analyzed and used to verify that the text was generated by the particular AI model.
The idea of "watermarking" AI-generated text was presented by Google as early as last year, when the company introduced its SynthID identifiers. And the online community is fairly convinced that Anthropic's text watermarks are based on SynthID technology or on a closely related solution.
Anthropic itself refers to a "statistical pattern" that allows the watermark to be detected even in long pieces of text that have subsequently been edited manually.
Another industry giant, OpenAI, has already confirmed outright that it uses Google's SynthID to watermark text generated by ChatGPT. And, naturally, Google's own Gemini AI models already use the same technology.
So going forward, AI-generated text that has not been radically edited by a human after generation can be identified with a fairly high degree of confidence - even if the quality of AI-generated text continues to improve beyond today's level.