top of page

In Defence of the Typo: AI-generated text to include Watermarks



Spotting text written by AI used to be easy, but very quickly the ease of looking for an em dash or keywords became harder with more complex tells like certain sentence structures, usually not X but Y, or examples given in threes, cropping up. This meant it became harder to decipher AI text from original text, and nearly impossible to prove with certainty. Luckily, this will now become much easier due to EU regulations.


There is about to be a watermark for those trying to detect AI that both humans and bots can use. The EU AI Transparency Act has introduced a code of practice that obliges all AI text to be watermarked. Gemini has already integrated those features (though removable) while Claude and ChatGPT are incorporating watermarking into their next update as of this month.

It will be automatically integrated into their future versions of the software and is planned to be backdated to current and older models as well. Anthropic (Claude.ai) have said that this watermark will apply worldwide, not just in Europe.


Article 50 of the AI Act attempts to tackle the uncertainty of what AI-written text is. It mandates that providers and deployers of AI systems need to label their output. The main objective of this is to tackle risks of deception and manipulation caused by AI, leading to increased integrity in the information ecosystem. This works into their wider frameworks, like the Digital Services Act (DSA), initiatives like the European Democracy Shield and other rules like those for high-risk AI systems or general-purpose AI models. They pertain to marking and detection of AI-generated content and the labelling of deepfakes and certain AI-generated publications.


Perhaps the most interesting element of this change is how it is proposed to work. Anthropic will not use any metadata or hidden characters in their output, but the watermark will be the very choice of words in the text. Basically, all possible word combinations have a statistical likelihood of having been used by any given AI model. For example, the sentence: ‘’Come in, sit on the couch.’’ I could have said couch, or sofa, or settee, etc. Each word has a token code to identify it.


Each AI model will have its own tendencies and word preferences that will be easily spotted by AI detectors. The method will provide a mathematical pattern that a detector can pick up. This means that it will not change the quality of any text written by Claude. However, it also means that copying and pasting it to and from your notes app or editing the image will not remove the watermark. This is easier to find in longer chunks of text. So, it has yet to be seen if AI-generated titles or slogans can be detected.

Within days, many tools to remove this watermark appeared on GitHub. The aim of these tools is to remove any special characters and change word choices to avoid detection. 


The watermark removers work by having a second AI reword your text to evade these hidden patterns by selecting random synonyms. However, it is hard to imagine this could read naturally. Results of evasion are not guaranteed. It also begs the question about the damaging quality of your prose by asking a cheaper AI to reword some text. It may result in the text losing meaning or perhaps even become a tell-tale sign of a watermark remover over time.


One may have to ask themselves what the point is of going to this extent to remove a watermark? If one needs to go to such lengths to hide the fact they are using AI, then perhaps they should not be using it.  Those who use AI to proofread or translate may now need to find another way to do that. Any text that may have even been tweaked with AI may now be susceptible to the watermark.


An interesting element to this is that as LLM’s improve with each update they mirror human writing more accurately. And in the same vein, the more human society is surrounded in AI text and videos, the more we subconsciously mirror the style, word choice and intonations back to the bot. The Guardian phrased it nicely: “The problem is that not only does AI train on human writing, but humans are stylistically influenced by AI, the interplay creating a kind of linguistic hall of mirrors.’’Current AI detectors are not without fault, so the new watermarking technique may prove to be a blessing. It may encourage people to start writing and trusting their own literacy levels. Typos and incorrect grammar are not fantastic to find in an email but have become a symbol that at least someone took the time to sit and write to you, rather than have an AI bot respond to them.

It will be interesting to see how this pans out, will we see a decrease in the use of AI text? Maybe there will be a new popularity around AI watermark removal tools, and a decrease in quality of work? Will human and LLM prose style become so enmeshed that it will be indistinguishable? Or perhaps the spotlight shining on those who use it will make it more socially acceptable in some jobs…


With the rise of watermarked AI, the effort of dodging detection may end up costing more than just sitting down and writing. In my view I hope that by trying to make AI text detectable, the EU will make human text valuable again, typos, awkward phrasing, and all.


Image: Unsplash/Emiliano Vittoriosi

No image changes made.

Comments


bottom of page