Want to bypass Turnitin in 2026? Grab the free prompt pack.
Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.
Send me the free prompts →Key Takeaways
- 30% / 20% / 17% detection on pristine watermarked text for KGW, SynthID-Text and Unigram, across 846 valid paraphrase runs (Tamim and Khan, 2026)
- 5.4% of clean human-written texts were flagged as watermarked by SynthID-Text in the same evaluation (Tamim and Khan, 2026)
- 98.3% to 100% of watermarks that were detected at all were gone after a single paraphrase pass (Tamim and Khan, 2026)
- 98.4% detection with zero false positives is the figure everyone quotes. It is a 2023 laboratory result at roughly 200 tokens (Kirchenbauer et al., via IEEE Spectrum)
- below 50% detection accuracy on short replies, against up to 95% in best-case conditions (IEEE Spectrum, 2026)
- above 90% scrubbing success against SynthID-Text using nothing more sophisticated than a baseline paraphraser (ETH Zurich SRI Lab, 2024)
- 3 points of correctness lost on code, while on prose the effect did not exceed changing the sampling seed (Nemecek et al., 2026)
- zero public tools can verify a text watermark for any provider, Claude and Gemini included, as of mid-September 2026
Anyone searching for AI text watermarking statistics runs into the same number within about thirty seconds: 98.4 percent detection, zero false positives. It appears in explainer after explainer as though it describes the watermark running in production today. It does not. It is a 2023 laboratory result, and the only independent evaluation published in 2026 found something considerably less reassuring. This page collects every figure that has actually been published, dates each one, and names the paper it came from. Where a number reached us through secondary coverage rather than the source paper, it says so, because the gap between those two things is exactly where the confusion in this topic lives. If you are here because a tool or an institution has made a claim about your writing, the numbers below sit alongside what detector false positive rates actually look like by tool.
1What detection rates do AI text watermarks actually achieve?
Between 17 and 30 percent on pristine, unattacked text, according to the only independent 2026 evaluation. The widely quoted 98.4 percent figure comes from a 2023 paper measuring ideal conditions at roughly 200 tokens, and reporting on the same research puts real-world accuracy at up to 95 percent in the best case and below 50 percent on short replies.
| Metric | Value | Source |
|---|---|---|
| KGW, detection on pristine watermarked text | 30% | Tamim and Khan (2026) |
| SynthID-Text, detection on pristine watermarked text | 20% | Tamim and Khan (2026) |
| Unigram, detection on pristine watermarked text | 17% | Tamim and Khan (2026) |
| Kirchenbauer et al., ideal conditions at ~200 tokens | 98.4% | 2023 result, via IEEE Spectrum (2026) |
| Best-case accuracy in reporting on current schemes | up to 95% | IEEE Spectrum (2026) |
| Accuracy on short replies | below 50% | IEEE Spectrum (2026) |
| Primary sources: arXiv:2607.16010 and IEEE Spectrum | ||
The three top bars are false negative rates in disguise. A 30 percent detection rate means 70 percent of genuinely watermarked text passed as unmarked, which puts these schemes in the same territory as the false negative rates already documented across conventional AI detectors. The gap between the top three bars and the bottom one is not a story about watermarking degrading over time. It is the difference between a controlled benchmark at a fixed token length and an evaluation that ran the schemes the way a third party would have to. It also explains why every tool currently advertising Claude watermark detection is looking for the wrong signal: there is no published detector for them to be wrapping.
2How often does a watermark detector flag human writing?
SynthID-Text flagged 5.4 percent of clean human-written texts as watermarked in the 2026 evaluation. That rate sits on top of an already low true detection rate, which is the worst combination available: the detector misses most real watermarks while still accusing roughly one human text in twenty.
A one-in-twenty error rate is not an abstraction once it is applied at the scale of a university intake or a publishing platform’s submission queue. It lands hardest on the same people conventional detectors already mistreat: writers with formal, structured, low-variance prose. That pattern is well documented for neurodivergent writers whose natural style depresses perplexity scores and for second-language writers, where the bias has been measured repeatedly. Watermarking was supposed to remove that class of error by replacing statistical guesswork with a cryptographic signal. On these numbers it has inherited the problem instead.
3Does paraphrasing remove a text watermark?
Yes, almost completely. Of the watermarks that were detected in the first place, a single paraphrase pass removed 100 percent for KGW and Unigram and 98.3 percent for SynthID-Text. A separate probe of SynthID-Text put scrubbing success above 90 percent using a baseline paraphraser.
| Metric | Value | Source |
|---|---|---|
| KGW, conditional removal after one paraphrase pass | 100% | Tamim and Khan (2026) |
| Unigram, conditional removal after one paraphrase pass | 100% | Tamim and Khan (2026) |
| SynthID-Text, conditional removal after one paraphrase pass | 98.3% | Tamim and Khan (2026) |
| SynthID-Text, scrubbing success with a baseline paraphraser | above 90% | ETH Zurich SRI Lab (2024) |
| SynthID-Text, spoofing success at default parameters | 4% | ETH Zurich SRI Lab (2024) |
| SynthID-Text, spoofing success at a 90,000 query budget | 15% | ETH Zurich SRI Lab (2024) |
| Words that must change to strip a watermark from a long response | roughly 25% | IEEE Spectrum (2026) |
| Primary sources: arXiv:2607.16010, ETH Zurich SRI Lab, IEEE Spectrum | ||
Note the word conditional. These percentages are calculated only across texts the detector caught to begin with, which was already a minority. The two findings compound rather than average: most watermarks were never detected, and nearly all of the ones that were did not survive being rewritten once. There is a useful asymmetry buried in the IEEE figure, though. If stripping a mark from a long response requires changing about a quarter of its words, then the mark is genuinely robust to light editing, which is why the tools selling themselves as Claude watermark removers have so little to actually do. Robust to a typo pass and defeated by a full rewrite is a coherent position, and it maps closely onto the bypass rates measured after humanization across conventional detectors.
4Which AI models watermark text, and can anyone verify it?
Anthropic and Google watermark text output today. OpenAI has not shipped a text watermark and says it plans to. No provider has released a public verification tool, so as of mid-September 2026 there is no way for a third party to check a text for any of these marks.
| Provider | Status | Source |
|---|---|---|
| Anthropic (Claude) | Watermarking text since 2 August 2026, files also signed with C2PA metadata, third-party detection tooling described as forthcoming | Anthropic |
| Google (Gemini) | Watermarks text output via SynthID-Text | IEEE Spectrum (2026) |
| OpenAI | No text watermark shipped, states it plans to | IEEE Spectrum (2026) |
| Public verification tools available to third parties | none | Independent five-method test (2026) |
| Primary sources: Anthropic, IEEE Spectrum, five-method test | ||
The verification gap is the practical headline. A watermark that only its issuer can read is not evidence available to a teacher, an editor or a hiring manager, which is a different situation from the one facing institutions already running commercial detectors they can query directly. Until a detection endpoint exists, any product claiming to check text for a Claude or Gemini watermark is inferring rather than reading.
5What does watermarking cost in output quality?
Almost nothing on prose. One 2026 paper found the measured effect of the watermark on prose did not exceed the effect of changing the sampling seed. On code it cost three points of correctness on one model and was below measurement on the other, while detection stayed near chance.
This is the one genuinely good result in the literature, and it cuts against the most common objection to watermarking. The worry was that biasing token selection would degrade writing. On prose it does not, at least not measurably. The problem is the other half of the same sentence: quality held up and detection did not. Short texts remain the weakest case across every approach in this field, the same structural reason short essays behave so badly under conventional AI detectors.
6Why this evidence does not meet a courtroom standard
The authors of the 2026 evaluation state plainly that these configurations, as tested, do not meet the evidentiary bar that courts require, and argue the evidence fails Daubert admissibility. Their recommendation is serious skepticism until methods demonstrate stable error rates under realistic adversarial conditions.
| Metric | Value | Source |
|---|---|---|
| Schemes evaluated | KGW, Unigram, SynthID-Text (via MarkLLM) | Tamim and Khan (2026) |
| Valid paraphrase runs | 846 | Tamim and Khan (2026) |
| Prompts per method | 15 | Tamim and Khan (2026) |
| Stated conclusion on admissibility | Fails the evidentiary bar courts require | Tamim and Khan (2026) |
| ETH Zurich probe scale | 30,000 black-box queries, 2,000 spoofing attempts, texts averaging 1,000 tokens | ETH Zurich SRI Lab (2024) |
| Primary sources: arXiv:2607.16010, ETH Zurich SRI Lab | ||
An error rate that is unstable under adversarial conditions is the specific thing a Daubert challenge is designed to surface, and it is the reason detector output has fared poorly whenever it has been tested properly, a pattern visible across the student cases that have reached a documented outcome. For anyone weighing whether a watermark claim could support a finding of misconduct, the relevant comparison is not the 98.4 percent headline but the rules of evidence, and those are already clear about what detector output can and cannot be used to prove.
7Watermark reliability checker
Select a scheme and a condition to see the published figures for that combination. Every value shown is drawn from the sources cited on this page. Nothing is modelled, estimated or extrapolated.
Published reliability by scheme and condition
Figures from Tamim and Khan (2026) unless otherwise noted.
Source: Tamim and Khan, arXiv:2607.16010 (2026).
Methodology
Research date: 28 September 2026. Every figure on this page was traced to its originating paper or the publication that reported it, and each is labelled with the year the measurement was made rather than the year it was written about.
- Source freshness. Four sources are from 2026, two are from 2024, and one measurement dates to 2023. The 2023 figure is retained only because it is the number most widely quoted, and it is dated everywhere it appears here.
- Corrections applied during research. Two figures circulating in aggregator coverage did not survive checking. A widely repeated claim that a single rewrite drops every scheme below 30 percent true positive rate was presented in secondary coverage as new August 2026 research. It traces to WaterPark, published as Watermark under Fire at EMNLP 2025 Findings and submitted in November 2024, and the figure does not appear in that paper’s abstract. It is therefore excluded from the tables above and noted here instead. A second claim that paraphrasing had stripped SynthID-Text was asserted without citation and has been replaced with the ETH Zurich probe, which states its metric, its paraphraser and its sample size.
- Excluded. One community reimplementation reported 88 percent accuracy and a 97.76 percent ROC-AUC over 50 generations on an open 8B model. It is a single hobbyist test at n=50 and is not placed beside peer-reviewed figures, though it is recorded here for completeness and needs independent replication.
- Limitations. The 2026 evaluation covers three schemes implemented through MarkLLM, not the production configurations running inside Claude or Gemini, which are not publicly testable. No published figures exist for Anthropic’s deployed watermark because no detection endpoint has been released.
- Update schedule. Reviewed when any provider ships a public verification tool, and otherwise quarterly.
Frequently asked questions
Can anyone check whether my text has an AI watermark?
Not as of mid-September 2026. An independent test of five different methods found no public tool that can verify a text watermark for any provider, Claude and Gemini included. Anthropic has described third-party detection tooling as forthcoming. Any product currently claiming to detect a Claude watermark is inferring from other signals rather than reading the mark.
Does Claude watermark everything it writes?
Anthropic began watermarking text on 2 August 2026 and says all future Claude models will produce watermarked text, with files also signed using C2PA metadata. Older models were described as being updated during a transition period, so text generated before that date is a different case from text generated after it.
Will paraphrasing remove a watermark?
On the published evidence, nearly always. One paraphrase pass removed 100 percent of detected watermarks for KGW and Unigram and 98.3 percent for SynthID-Text. A separate probe put scrubbing success above 90 percent with a baseline paraphraser. Light editing is a different matter: stripping a mark from a long response appears to require changing roughly a quarter of its words.
Is watermarking more accurate than a normal AI detector?
Not on the 2026 numbers. Detection on pristine watermarked text ran between 17 and 30 percent depending on the scheme, with a 5.4 percent false positive rate on clean human text for SynthID-Text. That is broadly comparable to the territory conventional tools occupy, and worth reading alongside how much detectors disagree with each other on the same text.
Where does the 98.4 percent figure come from?
Kirchenbauer et al., in 2023, measuring responses of roughly 200 tokens under laboratory conditions, reporting 98.4 percent detection with zero false positives. It is widely quoted without its date. Reporting on the same body of research puts realistic accuracy at up to 95 percent in the best case and below 50 percent on short replies.
Does watermarking make the writing worse?
Barely, on prose. A 2026 paper found the measured effect on prose did not exceed that of changing the sampling seed. On code it cost three points of correctness on one model and was below measurement on the other. The authors’ concern was detectability, which remained near chance, rather than quality.
Could a watermark be used to accuse me of using AI?
The authors of the 2026 evaluation argue it should not be, stating that these configurations as tested do not meet the evidentiary bar that courts require and fail Daubert admissibility. If you are facing an accusation now, the immediate steps matter more than the research, and there is a checklist for the first 24 hours after an AI accusation.
Sources
- Tamim, Saifur Rahman and Khan, Amir Labib. “AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation.” arXiv:2607.16010. arxiv.org. Submitted 17 July 2026. Accessed 28 September 2026.
- Jovanović, Nikola; Gloaguen, Thibaud and Vechev, Martin. “Probing Google DeepMind’s SynthID-Text Watermark.” SRI Lab, ETH Zurich. sri.inf.ethz.ch. Published 20 December 2024. Accessed 28 September 2026.
- Nemecek, Alexander; Chaudhary, Vipin and Ayday, Erman. “Watermarks Without Verification: AI Text Watermarking After the EU AI Act.” arXiv:2609.09604. arxiv.org. Submitted 9 September 2026. Accessed 28 September 2026.
- IEEE Spectrum. “AI Models Are Watermarking Text, Will You Notice?” spectrum.ieee.org. Published 10 September 2026. Accessed 28 September 2026.
- Anthropic. “How Claude’s text watermarking works.” anthropic.com. Published 14 August 2026. Accessed 28 September 2026.
- “Watermark under Fire: A Robustness Evaluation of LLM Watermarking” (the WaterPark benchmark). arXiv:2411.13425, EMNLP 2025 Findings. arxiv.org. Submitted 20 November 2024. Accessed 28 September 2026.
- “Robustness Assessment and Enhancement of Text Watermarking for Google’s SynthID.” arXiv:2508.20228. arxiv.org. Accessed 28 September 2026.
- “How to Check if Text Has an AI Watermark (I Tested 5 Methods).” medium.com. Published 16 September 2026. Accessed 28 September 2026.
Last updated 28 September 2026. All figures verified against their originating sources on that date.
