AI Text Watermarking: Every Published Number on Detection, False Positives and Paraphrase Survival (2026)

Published:

Updated:

Detection Drama · Free Download

Want to bypass Turnitin in 2026? Grab the free prompt pack.

Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.

Send me the free prompts →
Free · No credit card · Straight to your inbox
17% to 30%
Detection rates the three tested watermarking schemes achieved on their own pristine, unattacked watermarked text. Before anyone paraphrased anything, the best performer found fewer than one watermark in three.
Source: Tamim and Khan, arXiv:2607.16010 (2026)

Key Takeaways

  • 30% / 20% / 17% detection on pristine watermarked text for KGW, SynthID-Text and Unigram, across 846 valid paraphrase runs (Tamim and Khan, 2026)
  • 5.4% of clean human-written texts were flagged as watermarked by SynthID-Text in the same evaluation (Tamim and Khan, 2026)
  • 98.3% to 100% of watermarks that were detected at all were gone after a single paraphrase pass (Tamim and Khan, 2026)
  • 98.4% detection with zero false positives is the figure everyone quotes. It is a 2023 laboratory result at roughly 200 tokens (Kirchenbauer et al., via IEEE Spectrum)
  • below 50% detection accuracy on short replies, against up to 95% in best-case conditions (IEEE Spectrum, 2026)
  • above 90% scrubbing success against SynthID-Text using nothing more sophisticated than a baseline paraphraser (ETH Zurich SRI Lab, 2024)
  • 3 points of correctness lost on code, while on prose the effect did not exceed changing the sampling seed (Nemecek et al., 2026)
  • zero public tools can verify a text watermark for any provider, Claude and Gemini included, as of mid-September 2026

Anyone searching for AI text watermarking statistics runs into the same number within about thirty seconds: 98.4 percent detection, zero false positives. It appears in explainer after explainer as though it describes the watermark running in production today. It does not. It is a 2023 laboratory result, and the only independent evaluation published in 2026 found something considerably less reassuring. This page collects every figure that has actually been published, dates each one, and names the paper it came from. Where a number reached us through secondary coverage rather than the source paper, it says so, because the gap between those two things is exactly where the confusion in this topic lives. If you are here because a tool or an institution has made a claim about your writing, the numbers below sit alongside what detector false positive rates actually look like by tool.

1What detection rates do AI text watermarks actually achieve?

Between 17 and 30 percent on pristine, unattacked text, according to the only independent 2026 evaluation. The widely quoted 98.4 percent figure comes from a 2023 paper measuring ideal conditions at roughly 200 tokens, and reporting on the same research puts real-world accuracy at up to 95 percent in the best case and below 50 percent on short replies.

MetricValueSource
KGW, detection on pristine watermarked text30%Tamim and Khan (2026)
SynthID-Text, detection on pristine watermarked text20%Tamim and Khan (2026)
Unigram, detection on pristine watermarked text17%Tamim and Khan (2026)
Kirchenbauer et al., ideal conditions at ~200 tokens98.4%2023 result, via IEEE Spectrum (2026)
Best-case accuracy in reporting on current schemesup to 95%IEEE Spectrum (2026)
Accuracy on short repliesbelow 50%IEEE Spectrum (2026)
Primary sources: arXiv:2607.16010 and IEEE Spectrum
Detection on pristine watermarked text, before any attack
KGW
30%
SynthID-Text
20%
Unigram
17%
Kirchenbauer 2023 (lab)
98.4%
The bottom bar is a 2023 result under ideal conditions and is shown for contrast, not as a current measurement.

The three top bars are false negative rates in disguise. A 30 percent detection rate means 70 percent of genuinely watermarked text passed as unmarked, which puts these schemes in the same territory as the false negative rates already documented across conventional AI detectors. The gap between the top three bars and the bottom one is not a story about watermarking degrading over time. It is the difference between a controlled benchmark at a fixed token length and an evaluation that ran the schemes the way a third party would have to. It also explains why every tool currently advertising Claude watermark detection is looking for the wrong signal: there is no published detector for them to be wrapping.

Conceptual diagram of how a statistical AI text watermark biases token selection
A statistical text watermark is a bias in which words get picked, not a hidden character or a piece of metadata.

2How often does a watermark detector flag human writing?

SynthID-Text flagged 5.4 percent of clean human-written texts as watermarked in the 2026 evaluation. That rate sits on top of an already low true detection rate, which is the worst combination available: the detector misses most real watermarks while still accusing roughly one human text in twenty.

5.4%
of clean human-written texts were flagged as watermarked by SynthID-Text. For comparison, the headline 2023 figure that dominates coverage of this topic reported zero false positives.
Tamim and Khan, arXiv:2607.16010 (2026)

A one-in-twenty error rate is not an abstraction once it is applied at the scale of a university intake or a publishing platform’s submission queue. It lands hardest on the same people conventional detectors already mistreat: writers with formal, structured, low-variance prose. That pattern is well documented for neurodivergent writers whose natural style depresses perplexity scores and for second-language writers, where the bias has been measured repeatedly. Watermarking was supposed to remove that class of error by replacing statistical guesswork with a cryptographic signal. On these numbers it has inherited the problem instead.

0%
false positives in the Kirchenbauer result, at roughly 200 tokens under laboratory conditions in 2023. Quoting this figure without its year is the single most common error in current coverage of AI text watermarking.
Kirchenbauer et al., reported by IEEE Spectrum (2026)

3Does paraphrasing remove a text watermark?

Yes, almost completely. Of the watermarks that were detected in the first place, a single paraphrase pass removed 100 percent for KGW and Unigram and 98.3 percent for SynthID-Text. A separate probe of SynthID-Text put scrubbing success above 90 percent using a baseline paraphraser.

MetricValueSource
KGW, conditional removal after one paraphrase pass100%Tamim and Khan (2026)
Unigram, conditional removal after one paraphrase pass100%Tamim and Khan (2026)
SynthID-Text, conditional removal after one paraphrase pass98.3%Tamim and Khan (2026)
SynthID-Text, scrubbing success with a baseline paraphraserabove 90%ETH Zurich SRI Lab (2024)
SynthID-Text, spoofing success at default parameters4%ETH Zurich SRI Lab (2024)
SynthID-Text, spoofing success at a 90,000 query budget15%ETH Zurich SRI Lab (2024)
Words that must change to strip a watermark from a long responseroughly 25%IEEE Spectrum (2026)
Primary sources: arXiv:2607.16010, ETH Zurich SRI Lab, IEEE Spectrum
Share of detected watermarks removed by one paraphrase pass
KGW
100%
Unigram
100%
SynthID-Text
98.3%
Conditional removal, meaning the share of texts that were detected before the attack and were not detected after it.

Note the word conditional. These percentages are calculated only across texts the detector caught to begin with, which was already a minority. The two findings compound rather than average: most watermarks were never detected, and nearly all of the ones that were did not survive being rewritten once. There is a useful asymmetry buried in the IEEE figure, though. If stripping a mark from a long response requires changing about a quarter of its words, then the mark is genuinely robust to light editing, which is why the tools selling themselves as Claude watermark removers have so little to actually do. Robust to a typo pass and defeated by a full rewrite is a coherent position, and it maps closely onto the bypass rates measured after humanization across conventional detectors.

Illustration of an AI text watermark signal disappearing after the text is rewritten
Rewriting removed nearly every watermark that was detected in the first place.

4Which AI models watermark text, and can anyone verify it?

Anthropic and Google watermark text output today. OpenAI has not shipped a text watermark and says it plans to. No provider has released a public verification tool, so as of mid-September 2026 there is no way for a third party to check a text for any of these marks.

ProviderStatusSource
Anthropic (Claude)Watermarking text since 2 August 2026, files also signed with C2PA metadata, third-party detection tooling described as forthcomingAnthropic
Google (Gemini)Watermarks text output via SynthID-TextIEEE Spectrum (2026)
OpenAINo text watermark shipped, states it plans toIEEE Spectrum (2026)
Public verification tools available to third partiesnoneIndependent five-method test (2026)
Primary sources: Anthropic, IEEE Spectrum, five-method test

The verification gap is the practical headline. A watermark that only its issuer can read is not evidence available to a teacher, an editor or a hiring manager, which is a different situation from the one facing institutions already running commercial detectors they can query directly. Until a detection endpoint exists, any product claiming to check text for a Claude or Gemini watermark is inferring rather than reading.

5What does watermarking cost in output quality?

Almost nothing on prose. One 2026 paper found the measured effect of the watermark on prose did not exceed the effect of changing the sampling seed. On code it cost three points of correctness on one model and was below measurement on the other, while detection stayed near chance.

3 pts
of correctness lost on code for one tested model, below measurement on the other. On prose the watermark’s measured effect did not exceed that of changing the sampling seed, and the authors note detection remained near chance, which they frame as a limitation of detectability rather than quality.
Nemecek, Chaudhary and Ayday, arXiv:2609.09604 (2026)

This is the one genuinely good result in the literature, and it cuts against the most common objection to watermarking. The worry was that biasing token selection would degrade writing. On prose it does not, at least not measurably. The problem is the other half of the same sentence: quality held up and detection did not. Short texts remain the weakest case across every approach in this field, the same structural reason short essays behave so badly under conventional AI detectors.

6Why this evidence does not meet a courtroom standard

The authors of the 2026 evaluation state plainly that these configurations, as tested, do not meet the evidentiary bar that courts require, and argue the evidence fails Daubert admissibility. Their recommendation is serious skepticism until methods demonstrate stable error rates under realistic adversarial conditions.

MetricValueSource
Schemes evaluatedKGW, Unigram, SynthID-Text (via MarkLLM)Tamim and Khan (2026)
Valid paraphrase runs846Tamim and Khan (2026)
Prompts per method15Tamim and Khan (2026)
Stated conclusion on admissibilityFails the evidentiary bar courts requireTamim and Khan (2026)
ETH Zurich probe scale30,000 black-box queries, 2,000 spoofing attempts, texts averaging 1,000 tokensETH Zurich SRI Lab (2024)
Primary sources: arXiv:2607.16010, ETH Zurich SRI Lab

An error rate that is unstable under adversarial conditions is the specific thing a Daubert challenge is designed to surface, and it is the reason detector output has fared poorly whenever it has been tested properly, a pattern visible across the student cases that have reached a documented outcome. For anyone weighing whether a watermark claim could support a finding of misconduct, the relevant comparison is not the 98.4 percent headline but the rules of evidence, and those are already clear about what detector output can and cannot be used to prove.

7Watermark reliability checker

Select a scheme and a condition to see the published figures for that combination. Every value shown is drawn from the sources cited on this page. Nothing is modelled, estimated or extrapolated.

Published reliability by scheme and condition

Figures from Tamim and Khan (2026) unless otherwise noted.

Published detection rate20%
Watermarks missed80%
False positives on clean human text5.4%
Public tool to verify this yourselfNone
Detection missed 80% of genuinely watermarked text before any attack.

Source: Tamim and Khan, arXiv:2607.16010 (2026).

Methodology

Research date: 28 September 2026. Every figure on this page was traced to its originating paper or the publication that reported it, and each is labelled with the year the measurement was made rather than the year it was written about.

  • Source freshness. Four sources are from 2026, two are from 2024, and one measurement dates to 2023. The 2023 figure is retained only because it is the number most widely quoted, and it is dated everywhere it appears here.
  • Corrections applied during research. Two figures circulating in aggregator coverage did not survive checking. A widely repeated claim that a single rewrite drops every scheme below 30 percent true positive rate was presented in secondary coverage as new August 2026 research. It traces to WaterPark, published as Watermark under Fire at EMNLP 2025 Findings and submitted in November 2024, and the figure does not appear in that paper’s abstract. It is therefore excluded from the tables above and noted here instead. A second claim that paraphrasing had stripped SynthID-Text was asserted without citation and has been replaced with the ETH Zurich probe, which states its metric, its paraphraser and its sample size.
  • Excluded. One community reimplementation reported 88 percent accuracy and a 97.76 percent ROC-AUC over 50 generations on an open 8B model. It is a single hobbyist test at n=50 and is not placed beside peer-reviewed figures, though it is recorded here for completeness and needs independent replication.
  • Limitations. The 2026 evaluation covers three schemes implemented through MarkLLM, not the production configurations running inside Claude or Gemini, which are not publicly testable. No published figures exist for Anthropic’s deployed watermark because no detection endpoint has been released.
  • Update schedule. Reviewed when any provider ships a public verification tool, and otherwise quarterly.

Frequently asked questions

Can anyone check whether my text has an AI watermark?

Not as of mid-September 2026. An independent test of five different methods found no public tool that can verify a text watermark for any provider, Claude and Gemini included. Anthropic has described third-party detection tooling as forthcoming. Any product currently claiming to detect a Claude watermark is inferring from other signals rather than reading the mark.

Does Claude watermark everything it writes?

Anthropic began watermarking text on 2 August 2026 and says all future Claude models will produce watermarked text, with files also signed using C2PA metadata. Older models were described as being updated during a transition period, so text generated before that date is a different case from text generated after it.

Will paraphrasing remove a watermark?

On the published evidence, nearly always. One paraphrase pass removed 100 percent of detected watermarks for KGW and Unigram and 98.3 percent for SynthID-Text. A separate probe put scrubbing success above 90 percent with a baseline paraphraser. Light editing is a different matter: stripping a mark from a long response appears to require changing roughly a quarter of its words.

Is watermarking more accurate than a normal AI detector?

Not on the 2026 numbers. Detection on pristine watermarked text ran between 17 and 30 percent depending on the scheme, with a 5.4 percent false positive rate on clean human text for SynthID-Text. That is broadly comparable to the territory conventional tools occupy, and worth reading alongside how much detectors disagree with each other on the same text.

Where does the 98.4 percent figure come from?

Kirchenbauer et al., in 2023, measuring responses of roughly 200 tokens under laboratory conditions, reporting 98.4 percent detection with zero false positives. It is widely quoted without its date. Reporting on the same body of research puts realistic accuracy at up to 95 percent in the best case and below 50 percent on short replies.

Does watermarking make the writing worse?

Barely, on prose. A 2026 paper found the measured effect on prose did not exceed that of changing the sampling seed. On code it cost three points of correctness on one model and was below measurement on the other. The authors’ concern was detectability, which remained near chance, rather than quality.

Could a watermark be used to accuse me of using AI?

The authors of the 2026 evaluation argue it should not be, stating that these configurations as tested do not meet the evidentiary bar that courts require and fail Daubert admissibility. If you are facing an accusation now, the immediate steps matter more than the research, and there is a checklist for the first 24 hours after an AI accusation.

Sources

  1. Tamim, Saifur Rahman and Khan, Amir Labib. “AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation.” arXiv:2607.16010. arxiv.org. Submitted 17 July 2026. Accessed 28 September 2026.
  2. Jovanović, Nikola; Gloaguen, Thibaud and Vechev, Martin. “Probing Google DeepMind’s SynthID-Text Watermark.” SRI Lab, ETH Zurich. sri.inf.ethz.ch. Published 20 December 2024. Accessed 28 September 2026.
  3. Nemecek, Alexander; Chaudhary, Vipin and Ayday, Erman. “Watermarks Without Verification: AI Text Watermarking After the EU AI Act.” arXiv:2609.09604. arxiv.org. Submitted 9 September 2026. Accessed 28 September 2026.
  4. IEEE Spectrum. “AI Models Are Watermarking Text, Will You Notice?” spectrum.ieee.org. Published 10 September 2026. Accessed 28 September 2026.
  5. Anthropic. “How Claude’s text watermarking works.” anthropic.com. Published 14 August 2026. Accessed 28 September 2026.
  6. “Watermark under Fire: A Robustness Evaluation of LLM Watermarking” (the WaterPark benchmark). arXiv:2411.13425, EMNLP 2025 Findings. arxiv.org. Submitted 20 November 2024. Accessed 28 September 2026.
  7. “Robustness Assessment and Enhancement of Text Watermarking for Google’s SynthID.” arXiv:2508.20228. arxiv.org. Accessed 28 September 2026.
  8. “How to Check if Text Has an AI Watermark (I Tested 5 Methods).” medium.com. Published 16 September 2026. Accessed 28 September 2026.

Last updated 28 September 2026. All figures verified against their originating sources on that date.