AI text watermarks: OpenAI's textGrain, Claude's SynthID mark, and who gets to check them
What a text watermark result can establish, how much editing it survives, and why the error rates you can read so far were all measured by the vendor that holds the key. Then what that means for anyone who grades, writes or builds on these models.
On October 5 OpenAI announced that ChatGPT and Codex text in the European Union will carry an invisible watermark "over the coming weeks," and that API customers anywhere can switch the same watermark on today. The system is called textGrain. It adds no characters and no hidden symbols. It changes which words the model picks, slightly, in a pattern a secret key can later recognize.
Anthropic moved two months earlier. In August it announced that supported Claude models watermark their text on all of its products and in every country, with a version of the method Google DeepMind published in Nature in 2024. Both companies did it because of Article 50 of the EU AI Act, and both handled detection the same way: the detector is not public. Researchers and certain organizations can apply. A teacher, an editor or a writer who wants to check one document has no way in.
I read a watermark as a measurement instrument, and ask of it what I'd ask of any instrument: who holds it, who published its error rates, and who can repeat the test. For textGrain and Claude, the answers so far are the vendor, the vendor (or nobody), and nobody independent has published a result yet. That is not a new complaint. A September paper, Watermarks Without Verification, argues that this unverifiability, "rather than watermarking itself, is the substantive governance failure." One month later OpenAI has published numbers, which lets the question be asked more precisely.
What did OpenAI announce with the textGrain watermark?
Three things, in OpenAI's own announcement:
Starting today, API customers globally will be able to opt in to text
watermarking for select models. Text watermarking will remain off by
default in the API. Over the coming weeks, we will add an invisible
watermark to eligible ChatGPT and Codex text output in the European Union.The third is a detector: "Access will initially be limited to approved researchers and expert organizations." The help center names the first research partners: John Thickstun at Cornell, Martin Vechev at ETH Zurich and INSAIT, and the Kempelen Institute of Intelligent Technologies.
For API customers the switch is a dashboard setting, not a request parameter: Organization settings, Data controls, Text provenance, or the same in a project's settings, chosen per model. I found nothing per request in the API reference or changelog. For ChatGPT the scope is "all plans in the EU only," and OpenAI doesn't say how it decides who is in the EU.
OpenAI also published detection rates for its deployed watermark, which Anthropic hasn't done for its own:
| OpenAI's evaluation | Detected |
|---|---|
| 200-token passages, psychology-type questions, 1% false positive rate | about 80% |
| 400-token passages, same | about 95% |
| Math answers | "substantially lower," no figure given |
| 400-token passages, unedited (a separate test) | about 92% |
| Same, after replacing 10% of words with synonyms | 66% |
| Same, after replacing 25% of words with synonyms | 17% |
| 24 official EU languages, 1% false positive rate | 42.2% (Romanian) to 69.0% (Spanish) |
The language row comes from the help center, which says OpenAI then raised the watermark strength for languages below 60%. It doesn't give the test lengths, so that row can't be compared directly with the English ones.
A benchmark table for GPT-6 "Astra" with and without the watermark moves in both directions (GPQA Diamond 94.44% against 93.94%, Terminal-Bench 4.0 53.90% against 56.06%). OpenAI's reading is "we do not see meaningful performance differences." The technical report, written with researchers at UPenn and Yale, contains the math and no evaluation numbers. OpenAI says the report will be updated "in the coming weeks" and that it plans an open-source release, with no date.
Does Claude watermark its text?
Yes, and more broadly. Anthropic's explainer from August 14 and its support page say marks apply to output from supported models across the API, the Claude apps, Claude Code, Cowork and Claude Tag, through AWS, Google Cloud and Microsoft Foundry, "wherever Claude is offered, worldwide." Coverage is still being completed. The support page's model table includes models released before August, and the Bedrock rollout is due to finish by October 12. Neither page mentions a way to turn the watermark off.
Anthropic's reason for going global is one line:
We're applying watermarking globally at launch because we don't yet have
a durable way to scope it by region.The method is "a version of the SynthID-Text approach published by Google DeepMind." Anthropic answers practical questions OpenAI doesn't. Proofreading your text leaves "very little (if anything) for the watermark to attach to." A translation Claude produces is marked, "because in this case every word is chosen by Claude." On editing: "Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will." Detection is a "detection API in private preview," open to regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, EU civil society groups and companies with their own obligations under the Act.
Anthropic has published no detection rate at all: no length curve, no false positive rate, no editing test.
The August announcement drew the louder reaction. TechCrunch noted that Claude users objected they had supplied "the instructions, context, decisions" while Claude was just "the tool." In the threads I read this week, a recurring reply to OpenAI was relief that it stayed EU-only, and at least one reply suggested moving to Claude to avoid it. That doesn't help.
Which AI companies watermark text, and where?
As of October 6, from each company's pages where they exist:
| Provider | Text watermark | Scope | Detection access |
|---|---|---|---|
| OpenAI | textGrain (own method) | ChatGPT and Codex in the EU, rolling out; API opt-in worldwide | Application, approved researchers and organizations |
| Anthropic | Version of SynthID-Text | Supported models on all Claude surfaces and cloud partners, worldwide | Private-preview API for eligible organizations |
| SynthID-Text | Gemini app and web; a Google staffer confirmed in August that Gemini API text is marked, after first saying it wasn't | SynthID Detector portal, announced in 2025 with a waitlist for journalists and researchers | |
| Mistral | Not specified | Its Vibe Code legal page commits to marking and detection "by December 2, 2026" | Not specified |
| Meta, Microsoft | No deployment statement found | Both signed the provider section of the EU code | Not specified |
| Chinese providers | Visible labels and file metadata required since September 1, 2025; watermarks only "encouraged" | China | Not specified |
OpenAI already uses Google's SynthID for images and audio, but built its own scheme for text, for "more control than SynthID or TextSeal over the balance between watermark detectability and the variety of responses generated from the same prompt." TextSeal is Meta's open-source text watermark. I found no announcement of it running in a Meta product.
What does EU AI Act Article 50 require for generated text?
Article 50(2) is the binding part. Providers of systems that generate text, images, audio or video must ensure their outputs "are marked in a machine-readable format and detectable as artificially generated or manipulated," with solutions that are "effective, interoperable, robust and reliable as far as this is technically feasible." It doesn't prescribe a watermark. It applies from August 2, 2026, and a breach can cost up to EUR 15 million or 3% of worldwide annual turnover, whichever is higher (smaller companies get the lower of the two).
The Digital Omnibus, Regulation (EU) 2026/1744, gives systems that were on the market before August 2 until December 2, 2026. That fits OpenAI's "over the coming weeks," Mistral's December 2 date, and Anthropic marking new models at launch while still covering older ones. Some press reports say the Omnibus shortened a six-month grace period to three months. The published text gives four months.
The detail sits in two documents that carry less weight. The Code of Practice, final since June 10, is voluntary, and the Omnibus notes that such codes do not grant a presumption of conformity. Its provider section is signed by Anthropic, Google, Meta, Microsoft, Mistral, OpenAI and others. It normally asks for two layers, signed metadata plus an invisible watermark, but treats free-form text as unable to carry metadata, so a watermark alone is enough for it. It doesn't require watermarking text under 200 tokens.
On detection, the Code's default is generous: free of charge, with unlimited access for regulators, media, fact-checkers, independent researchers, educational and research institutions and civil society. Text is the exception:
Signatories may restrict access to detection mechanisms associated to
watermarking techniques for free-form text to the extent that they have a
lower level of reliability and robustness...Access may then be limited to verified expert users, and the Code says any restriction "will be limited in time," without naming a date. OpenAI cites the Code when it describes its gated access. Anthropic's list of eligible groups closely follows the Code's default list.
The Commission's Article 50 guidelines of July 20, also not binding, place source code outside the obligation, along with SQL, configuration files, infrastructure as code, comments that belong to the code, and messages exchanged only between machines, such as agent-to-agent traffic. They list "AI-generated translations of text" among the minor edits that don't need marking. OpenAI's help page attributes its code exclusion to the Code of Practice. It actually comes from the guidelines. Anthropic marks Claude's translations even though the guidelines would let it skip them.
How does an LLM text watermark work?
A language model writes one token at a time, sampling each from a probability distribution. A watermark tampers with that sampling, and only that, using a secret key.
The best-known early design is the green list from Kirchenbauer and colleagues at the University of Maryland (ICML 2023). Hash the previous token with a key, use the result to split the vocabulary into a "green" share and a "red" rest, and give green tokens a small bonus. Human text hits green about as often as chance predicts. Watermarked text hits it more often, and a z-test on the count tells them apart. With a 50/50 split and a moderate bonus, the paper detected 98.4% of 200-token generations at a false positive rate of 3 in 100,000. The bonus does shift the model's choices.
Scott Aaronson, then at OpenAI, described a scheme in 2022 that doesn't shift them on average. It picks each token with a keyed pseudorandom function, in a way that still matches the model's probabilities. The catch, as the textGrain report puts it: "at a fixed context and key, it always selects the same token." Same full context, same key, same settings, same answer.
Google DeepMind's SynthID-Text (Nature, October 2024) uses tournament sampling: draw many candidate tokens and run them through knockout rounds, each judged by a different keyed function. In a live test on about 20 million Gemini responses, user ratings for watermarked and plain answers differed by hundredths of a percent, not a statistically significant gap. On Gemma 7B in a separate benchmark, it added 0.57% to generation time. This is what Anthropic adapted.
textGrain keeps Aaronson's statistics without the identical answers. It sets an "entropy budget," an ideal bound on how much next-token sampling randomness the watermark takes on average, and couples the keyed values to the model's choices through an optimal-transport step. Averaged over keys, each next-token distribution stays the same. The detector "requires only the generated text and the secret key."
All of these share one limit. A watermark lives in choices the model could have made differently. When the next token is close to certain, there is nothing to tilt. OpenAI says math detects "substantially lower" because there is "less flexibility in word choice." Anthropic says where "an exact output is required... the watermark isn't applied," and code "has generally less watermarking," mostly in comments. Two inferences of mine follow from the mechanism, neither tested in public. The format alone shouldn't decide it: a JSON field holding a paragraph of prose has as much room as the paragraph. And text written in a controlled language built to limit word choice, like the ASD-STE100 standard Karpathy recommended for readable model output, should carry less signal.
The one independent number on code I found is in the September paper. Its authors ran Google's open-source SynthID implementation on two open-weight models. On prose, the watermark's effect on quality was no larger than changing the random seed. On code, correctness dropped three points on one model and was too small to measure on the other, "while detection remains near chance." That tests the public implementation, not Claude or textGrain, and it is still more than either vendor has published about detecting watermarks in code.
How is a text watermark detected, and who can run the detector?
The detector is a keyed statistics test. It recomputes the keyed values for the text, scores them (textGrain scores only the first occurrence of each distinct context window, so repeated phrases don't count twice) and asks how likely that score is for text unrelated to the key. In theory the false positive rate is set in advance rather than estimated from a benchmark, which is a real difference from style-based AI detectors. In practice the textGrain report adds that a deployed key needs "empirical calibration checks."
Evidence grows with length and entropy. Kirchenbauer's green list reached 98% at 200 tokens on a base model. Kuditipudi and colleagues (TMLR 2024) detected from as few as 35 tokens on base models, but on an instruction-tuned model with short answers only about a quarter of responses were detectable. OpenAI reports about 80% at 200 tokens. Short chat answers are the hard case for every scheme.
Only the key holder can compute the test, and each vendor has its own key. Anthropic notes its detector can't tell whether text came from a different AI, "even if that other AI uses watermarking, it would have a different key." So the error rates in OpenAI's table are the vendor grading its own instrument. Removal tools can't be validated against the real detectors. And the promise that a watermark does not identify the user is a claim only the key holder can check.
There is a real argument for gating, and OpenAI makes it: "Given the risk of missed watermarks and false positives, we are not making it publicly available at launch." A public detector is also an oracle. An attacker can query it until an edit just passes, or use it to confirm a forgery. The stealing attacks below need only model output, so gating doesn't stop them, but it denies the attacker a way to check the result. But a closed detector doesn't have to be the only design. Fairoze and colleagues describe a publicly detectable watermark whose detection algorithm "contains no secret information." A University of California patent application covers a signature embedded in text that can be verified with a public key. Keeping detection private is a choice, and a partly defensible one.
For anyone outside an approved institution, then, any website offering to check text for "the ChatGPT watermark" is running something else, usually a style classifier. Those have a weak record. OpenAI's own, launched in January 2023, caught 26% of AI text and flagged 9% of human text before it was withdrawn on July 20 that year "due to its low rate of accuracy." A study in Patterns ran seven detectors on TOEFL essays by non-native English speakers and found an average false positive rate of 61%. A watermark avoids that kind of style bias, but only for whoever is allowed to run it.
How do people remove or forge AI text watermarks?
OpenAI listed the routes in August 2024, when it held its earlier watermark back: "translation systems, rewording with another generative model, or asking the model to insert a special character in between every word and then deleting that character - making it trivial to circumvention by bad actors." For textGrain, its 2026 materials confirm the first two weaken detection. The third hasn't been tested in public.
Paraphrasing is the best studied. A dedicated paraphrase model, DIPPER, cut green-list detection from 100% to 57.2% (NeurIPS 2023). Paraphrasing repeatedly drove it below 20% after five rounds, and human raters still scored the output well. OpenAI's 25% synonym row is the same effect, though 17% is not zero. Length works the other way: Kirchenbauer's group found that text heavily paraphrased by humans was detectable after about 800 tokens on average, at a false positive rate of 1 in 100,000. A long essay leaks more than a short answer. Zhang and colleagues proved (ICML 2024, "Watermarks in the Sand") that strong watermarking is impossible if an attacker can judge quality and make small edits that, repeated long enough, can reach any comparable text. Their generic attack removed three published schemes with minor quality loss.
Translation is the second route: ask for the answer in another language and translate it back with a tool that doesn't watermark. On the schemes tested at ACL 2024, this pushed detection close to chance. How much survives depends on the design, and nobody has published a translation test of textGrain or Claude's mark. OpenAI and Google's SynthID documentation both list translation as a weakness. The translator has to be unmarked, so Claude won't do.
The insert-and-delete trick asks the model to put an emoji or symbol after every word, then deletes them. The keyed values were computed with the symbols in context, so after they are gone the pattern no longer lines up. Kirchenbauer's paper credits it to Riley Goodside.
Rewriting with an open-weight model that has no watermark is the route people talk about most. The SynthID paper names the gap itself: enforcing watermarking on open models "deployed in a decentralized manner" is difficult. Text that started on one never carried a mark.
Forging is the dangerous direction, because it means making a person's own writing test positive. Jovanović, Staab and Vechev at ETH Zurich estimated (ICML 2024) that for under $50 in API queries an attacker could learn enough about several earlier schemes to both forge and remove them, with over 80% success. The same lab probed the open SynthID-Text implementation. Forging was much harder, about 4%, rising to about 15% with 90,000 queries. Paraphrasing still scrubbed it over 90% of the time. Nobody has tested textGrain this way in public. Vechev is now one of OpenAI's evaluation partners.
A removal market appeared within weeks of Anthropic's announcement, with sites offering "Claude watermark removers." Without access to a vendor detector, none of them can show that their tool works. The "invisible character cleaners" left over from 2025 don't touch any of this. Back then, people found narrow no-break spaces in output from OpenAI's o3 and o4-mini and called them a watermark. The startup that reported it relayed OpenAI's reply that they were "a quirk of large-scale reinforcement learning." Today's watermarks have no characters to strip: OpenAI's "does not add hidden characters, invisible spaces, or unusual punctuation," and Anthropic says "Nothing is added to the text and there are no hidden characters."
What can an AI text watermark prove, and what can't it?
OpenAI's list of what a watermark doesn't tell you is the most useful part of its announcement. It does not measure human contribution, establish ownership or responsibility, or check accuracy. It "does not associate a person, organization, account, prompt, or conversation with the text." And:
The absence of a detected watermark does not prove human authorship.Anthropic adds that its detector "cannot distinguish 'Claude wrote this' from 'Claude heavily edited this.'" A positive result is statistical evidence that a model chose many of the words. Given a nonzero false positive rate and the forging research, it falls short of proof even of that. It says nothing about who had the ideas or whether using the model was allowed.
Then multiply by volume. OpenAI made this point against itself in 2024: "applying it to large volumes of text would lead to a large number of total false positives." It raised a fairness concern too: watermarking "could stigmatize use of AI as a useful writing tool for non-native English speakers." The guidelines exempt grammar fixes and light polishing, but a model that rewrites a paragraph for fluency is choosing words, and the mark can't tell why.
The law asks for marks that are "interoperable" and "robust." Each vendor uses its own key and its own detector, and OpenAI's own test shows detection at 17% after a quarter of the words change. Whether that meets "as far as this is technically feasible" is a question for the regulators, and the Code's text exception suggests they have already accepted the gap for now.
Who has filed patents on LLM text watermarking?
The filings I could verify are mostly about the sampling methods above. DeepMind's US20240320529A1, priority March 2023 and now assigned to GDM Holding, covers multi-stage keyed watermarking of each token, with dependent claims describing the tournament used in SynthID-Text. The University of Maryland group behind the green list filed US20250238634A1, priority January 2024, whose first claim is broad: modifying a model's sampling with pseudorandom scores assigned to words. An application assigned to UC San Diego (the search listing names Northeastern University), US20260113199A1, signs the first tokens with a key pair and embeds the signature in later ones, so it can be verified with a public key. Baidu's CN116340909A is the outlier: it encodes user information in the visual attributes of each character to trace screenshots. That is exactly what OpenAI and Anthropic say their watermarks don't do.
I found no published statistical text-watermark patent from OpenAI or Anthropic. That proves little. US applications are normally published about 18 months after their priority date, so recent filings may not be visible yet.
What should teachers, writers and developers do about AI text watermarks?
If you grade or edit other people's writing, you can't check text yourself today. Your institution can apply to OpenAI and to Anthropic, separately, and each detector only sees its own vendor's mark. A service that claims to check for the ChatGPT or Claude watermark is almost certainly guessing from style. A negative result proves nothing, and a positive one is evidence of model involvement, not of cheating.
If you write with a model, expect Claude output to be marked wherever you are, and ChatGPT output to be marked once the rollout reaches you, if you are in the EU. Proofreading your own text leaves little to mark. Asking a model to rewrite whole paragraphs leaves a lot. If a publisher or client distinguishes AI-assisted from AI-written work, keep your drafts, because the watermark can't make that distinction.
If you build on the APIs, the defaults differ. OpenAI's watermark is off unless someone in your organization turns it on, per project and per model. Claude's is on, and no opt-out is documented. Code, configuration and machine-to-machine messages fall outside the obligation as the Commission's guidelines read it, which is a legal line, not a technical guarantee that they carry no mark. Customer-facing prose your product generates most likely falls inside it. Anthropic's advice to builders is direct: "you should independently assess what Article 50 requires of your products and services." If you publish AI text on matters of public interest, Article 50(4) adds a disclosure duty unless the text went through human review or editorial control and someone holds editorial responsibility for it.
The next signs to watch are whether the textGrain release lets outsiders reproduce OpenAI's curves on their own models, whether Anthropic publishes an error rate, and whether the Code's restriction on text detectors is ever lifted.
My read is that there are two separate failures here, and only one of them can be fixed by a decision. The signal is fragile. That comes from the method, and it may never fully go away. The instrument is closed. That is a choice, defensible at launch and covered by the Code, and it can be reversed. Until it is, a text watermark is a measurement the vendor takes and reports on request: real evidence for a regulator, and very little for anyone who can't get approved.