Does WriteHuman Bypass GPTZero? Its Own Data Says It Varies (2026)

Published:

Updated:

Disclosure. Detection Drama earns affiliate commission from WriteHuman. This page credits WriteHuman for publishing the only reproducible dataset in this market, and also reports that its own data undercuts the number it advertised, that it ranks itself first in its own benchmark, and that GPTZero’s own test of it failed. No link to WriteHuman below is an affiliate link; outbound links are rel=nofollow citations. The promoted call-to-action above is a separate StealthWriter affiliate placement carrying its own disclosure.

Does WriteHuman bypass GPTZero? Its own published data says 61% one month and 88% the next — and it is the only company in this market that lets you check.

Detection Drama · 2026 Verdict

The only tool that bypassed Pangram and Turnitin in 2026.

Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.

Try StealthWriter →
Free plan to start · No credit card · Paid from $20/mo
Or get the free prompt pack first →
Affiliate link — I earn a commission if you upgrade. I only point at tools I actually run.

Key Takeaways

  • WriteHuman publishes the only reproducible benchmark in this market — every input, output, detector verdict and the scoring code, on GitHub under an open licence. We downloaded the raw files and checked the numbers ourselves.
  • Those files show its GPTZero bypass rate at 78.2% in June, 60.8% in July, 88.4% in August and 71.8% in September 2026 — a 27.6-point swing in four months, same tool, same method.
  • Its own August blog post headlined the 88.4% as the best result of any tool tested. Its own September data, from the same repository, is 71.8%.
  • GPTZero tested WriteHuman itself on 30 July 2026 and flagged the output at 84% AI paraphrasing — a fail. July is also the worst month in WriteHuman’s own data. The two sources agree.
  • WriteHuman runs the benchmark it ranks first in, in all four cycles. The July post disclosed that conflict; the August post did not.
  • In September, GPTZero was the hardest of the five detectors for it — below Copyleaks at 86.7% and Originality.ai at 91.4%. That inverts the usual pattern.
  • The category data matters more than the headline: academic essays 99.97%, but business emails 52.9%. What you are writing changes the answer more than which tool you buy.

By Vlad Ivanov — publisher of Detection Drama, tracking AI-detector and humanizer behaviour across 245 published tests and teardowns. Last updated: August 2026

Does WriteHuman bypass GPTZero - 2026 evidence review
Detection Drama · 2026 Verdict

The only tool that bypassed Pangram and Turnitin in 2026.

Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.

Try StealthWriter →
Free plan to start · No credit card · Paid from $20/mo
Or get the free prompt pack first →
Affiliate link — I earn a commission if you upgrade. I only point at tools I actually run.

Does WriteHuman bypass GPTZero?

Almost every page in this series ends the same way: the vendor claims a number, nobody can reproduce it, and the honest answer is that nobody knows.

This one is different, and the difference is worth stating up front. WriteHuman publishes its raw data — every source passage, every humanized output, every detector verdict, and the scoring code — in a public repository under an open licence. We did not take its word for anything below. We downloaded the files.

What they show is not the number on the marketing page. It is something more useful: the answer moves a lot, month to month, on the same test.

What does the raw data actually say?

WriteHuman GPTZero bypass rate by month June to September 2026

WriteHuman operates HumanizerBench, a monthly benchmark scoring humanizers against five detectors — GPTZero, Winston AI, ZeroGPT, Copyleaks and Originality.ai. We pulled its per-tool data file on 3 September 2026. Here is WriteHuman against GPTZero, cycle by cycle:

CycleGPTZero bypassHardest detector that cycleComposite
June 202678.2%Originality.ai (46.3%)83.59
July 202660.8% — worstOriginality.ai (53.6%)73.07
August 202688.4% — bestOriginality.ai (66.5%)76.69
September 202671.8%GPTZero (71.8%)78.29

Source: data/humanizers/writehuman.json, HumanizerBench public repository, downloaded and parsed 3 September 2026. Figures are bypass rates — the share of outputs the detector rated human.

A 27.6-point swing in four months, on the same benchmark, the same five detectors, the same stated method. Nothing about the product changed that much. What that range really measures is how unstable this question is.

Which makes one thing worth flagging. WriteHuman’s own August 2026 write-up headlined the 88.4% figure as “the best result posted by any of the twelve tools tested.” That was true of that cycle. One month later, its own repository records 71.8%. The company published both. The blog post only celebrates one.

Can you check GPTZero yourself?

GPTZero is free to test - 10,000 characters with no account

Turnitin sells no individual licences and shows students no AI report, so nobody outside an institution can produce a genuine Turnitin screenshot. GPTZero removes that excuse entirely. Anyone can paste up to 10,000 characters into gptzero.me right now with no account, no card and no signup; a free account raises the per-scan limit to 150,000 characters with a 10,000-word monthly quota, and paid tiers start at $9.99 a month.

So when a humanizer claims to beat GPTZero and shows you nothing, that is a choice, not a constraint.

It also means you can check your own text. A genuine current GPTZero result shows a three-class classification — human, AI or mixed — a confidence category rather than a bare percentage, a probability breakdown across all three classes, and sentence-by-sentence highlighting. An image showing only “X% AI”, or one displaying perplexity and burstiness, is either fabricated or dates to roughly 2023.

What happened when GPTZero tested it?

On 30 July 2026, GPTZero published its own hands-on review of WriteHuman. It ran a single 800-word ChatGPT-generated sample through WriteHuman and then through four detectors.

DetectorVerdict on WriteHuman outputRead
ZeroGPT0% AI / 100% humanPassed
Copyleaks0% AI, labelled entirely humanPassed
Originality.ai63% confident humanMarginal pass
GPTZero84% AI paraphrasingFailed

Source: gptzero.me/news/writehuman-ai-review, Mehal Rashid, 30 July 2026. Published by GPTZero about a tool that competes with GPTZero’s interests; n=1.

GPTZero’s verdict line was “We are highly confident this text was originally AI, but rewritten by AI or human” — the Paraphraser Shield firing — and its conclusion was that WriteHuman “couldn’t bypass GPTZero as it claims.” It went on to argue that “only poor AI detectors fall for AI humanizers.”

Take that for exactly what it is: one sample, published by the detector under test, on a page arguing its own product is superior. On its own it would not be worth much.

Except the dates line up, and that changes things. GPTZero ran its test at the end of July 2026. WriteHuman’s own repository records the July cycle as its worst GPTZero month of the four, at 60.8%. Two parties with directly opposed commercial interests, using entirely different methods, independently place late July as a weak period for WriteHuman against GPTZero. Agreement between adversaries is the strongest signal available in this market, and it is rare enough to say so.

How much should you trust a benchmark its winner owns?

Both halves of that sentence are true and both matter.

The creditable half: HumanizerBench publishes samples.json, tests.json, detector-scores.json and a frozen scoring.js for each cycle, under MIT and CC BY 4.0 licences. The stated method is that every tool is “bought and paid for by us and run by hand on the most evasion-focused setting it advertises”, with “no affiliate arrangements and no vendor-supplied numbers.” The September cycle covers fourteen tools across seven writing categories. Nobody else in this market has published anything close. We were able to verify every figure in this article against those files, which is not something we have been able to do once in the previous forty-five.

The other half: WriteHuman ranks first in every cycle it has published — 83.59, 73.07, 76.69 and 78.29, ahead of Stealth Writer, HIX Bypass and Undetectable AI. A scoring formula weights bypass at 42%, meaning preservation at 32%, readability at 16% and consistency at 10%, with penalties deducted for quality failures; those weights were chosen by the company that wins under them. And the July write-up disclosed the conflict of interest while the August one — the one headlining WriteHuman’s own top placement — did not.

Our reading: use the raw files, treat the leaderboard ordering with the caution any self-scored ranking deserves, and note that publishing data which visibly undercuts your own marketing is not the behaviour of a benchmark designed purely to flatter.

Which detector is actually the hard one?

This is where the usual story breaks down, and it is worth knowing before you buy anything on the strength of a GPTZero number.

Across most of this series, GPTZero is the softest detector on record — the UChicago BFI paper found it missing roughly half of humanized text while Pangram missed almost none. For WriteHuman in June and July, that held: Originality.ai was the hardest of the five, at 46.3% and 53.6%.

By September it had inverted. GPTZero, at 71.8%, was the hardest of WriteHuman’s five detectors — below Winston AI at 84.1%, ZeroGPT at 77.9%, Copyleaks at 86.7% and Originality.ai at 91.4%. One tool, one cycle, so do not over-read it. But it is a direct, checkable counter-example to “GPTZero is the easy one”, and it came from data the tool’s own maker published.

What matters more than the tool you pick?

What you are writing. The September category breakdown for WriteHuman, from the same file:

Writing categoryBypass rate
Discussion-board post100%
Academic essay99.97%
Marketing copy99.3%
Blog post99.4%
News article89.2%
Application essay76.0%
Business email52.9%

Source: HumanizerBench September 2026 cycle, category breakdown for WriteHuman, downloaded 3 September 2026. Aggregated across all five detectors, not GPTZero alone.

Same product, same month, same detectors: near-certainty on one genre and roughly a coin flip on another. The spread between categories is wider than the spread between most tools. Any vendor headline that gives you one number for “bypass rate” is averaging across a range that looks like this.

And what does GPTZero’s own paper say about it?

Read that table carefully, because it is easy to misread. The column is headed “bypassing ability on naive AI detection methods” — it rates those nine tools against weak detectors, not against GPTZero. The caption states GPTZero’s own position outright: “These methods are all ineffective against GPTZero due to the four-tiered red teaming approach.” That paper is also a GPTZero preprint, not peer-reviewed, and GPTZero has an obvious interest in the conclusion.

GPTZero’s February 2026 paper names nine bypass services in an appendix table. WriteHuman is rated “Medium”, with Grubby AI and TwainGPT; Undetectable is the only “High”, and StealthGPT, Quillbot, StealthWriter, HIX and GPTinf are “Low”.

How reliable are the detector numbers on either side?

GPTZero’s own FAQ concedes: “Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation.” Its developers page simultaneously claims a “Paraphraser Shield” means “even if AI content has been altered to look more human-like, GPTZero can detect it.” Both statements are currently published.

The two leading detector vendors publish opposite numbers on the same question. GPTZero’s February 2026 paper reports 93.5% recall on 1,000 texts run through nine bypasser services, and puts Pangram at 49.7%. Pangram’s DAMAGE paper reports GPTZero at 60.04% on humanized text, and 34.53% at GPTZero’s own default threshold. Each benchmark shows its publisher winning, and neither is peer-reviewed.

One figure to handle carefully. “GPTZero only catches 46% of humanized text” circulates constantly in humanizer marketing. It comes from Russell, Karpinska and Iyyer and it is one cell of one table: the o1-Pro humanized condition, n=30, where the “humanizer” was a research prompt the authors wrote themselves rather than any commercial tool. In the same table GPTZero scored 100% on paraphrased text and 85.3% overall at a 0.7% false positive rate. We will not quote the 46% without those qualifiers.

On false positives, GPTZero was one of seven detectors in Liang et al. (2023), which reported a 61.22% average false positive rate across those seven on 91 TOEFL essays. No per-detector figure was ever published, so that number is not GPTZero’s. GPTZero now claims it has cut its own TOEFL false positive rate to 1.1% — a figure no third party has verified.

Methodology. This page reports no first-party humanizing test. On 3 September 2026 we fetched WriteHuman’s sitemap, read its July and August 2026 ranking posts, and then downloaded the HumanizerBench repository’s per-tool data file and README directly from raw.githubusercontent.com rather than quoting the blog figures — every number in the tables above was parsed from that JSON and is reproducible from it. We confirmed the repository exists, is public, carries MIT and CC BY 4.0 licences, and publishes samples, humanized outputs, detector verdicts and the scoring script per cycle. We read GPTZero’s 30 July 2026 review of WriteHuman at source and record its authorship and publisher. GPTZero’s appendix table, access limits and FAQ statements are quoted from GPTZero’s own documents. The UChicago BFI figures come from that paper’s Tables 3 and B.4. Our affiliate relationship with WriteHuman is disclosed at the top, and the findings adverse to it are stated in the same detail as the favourable ones.

What should you take from this?

WriteHuman did the thing this market never does: it published the evidence, including the parts that make it look worse. That deserves saying plainly, and it is the reason this page could be written from data rather than from claims.

The data says its GPTZero performance is real but unstable — between roughly 61% and 88% across four consecutive months — that GPTZero’s own adversarial test failed it during its weakest month, that GPTZero was the hardest of five detectors for it in the most recent cycle, and that the genre you are writing swings the outcome by nearly fifty points on its own.

None of that supports betting a submission on a number. It supports checking your own specific text, which GPTZero lets you do free, in about a minute, with no account — and keeping your drafting history either way.

If your real worry is being wrongly flagged on your own writing, keep your drafting history — it costs nothing and it survives a detector being wrong. It matters most for the writers detectors treat worst: ESL writers and AI detection and false positives for neurodivergent students. Check before you submit with the best pre-submission check and the Turnitin self-check, and strip the obvious tells by hand first via what to remove before using an AI humanizer.

For the product itself see our WriteHuman review, and the same question against Turnitin and Pangram. Also in this cluster: StealthGPT, Clarity Bubble, GPTHuman.

Get the free prompt pack instead

The manual rewriting prompts that lower AI signals without a subscription — and without betting a submission on a number nobody has reproduced. Free, no card.

Grab the free prompts at detectiondrama.com →

Frequently asked questions

Does WriteHuman bypass GPTZero?

Sometimes, and the honest answer is that it varies far more than any marketing number admits. WriteHuman’s own published benchmark data records its GPTZero bypass rate at 78.2% in June 2026, 60.8% in July, 88.4% in August and 71.8% in September. That is a 27.6-point range across four consecutive cycles of the same test. Any single figure quoted from that series — including the 88.4% the company headlined — is a point on a volatile line, not a property of the product.

Where do those numbers come from?

From WriteHuman itself, and unusually, you can check them. It operates HumanizerBench, which publishes every source passage, every humanized output, every detector verdict and the scoring script in a public GitHub repository under MIT and CC BY 4.0 licences. We downloaded the per-tool data file directly rather than quoting the blog posts, and the four figures above come from that file. No other vendor in this series has published anything comparable.

Is a vendor-run benchmark trustworthy?

Treat it as evidence with a stake attached. WriteHuman pays for the tools, runs them by hand, and states that there are no affiliate arrangements and no vendor-supplied numbers — all creditable. It also ranks itself first in every cycle published. The July 2026 write-up disclosed the conflict; the August one, which headlined WriteHuman’s top placement, did not. The raw data is the reason to take it seriously; the ownership is the reason to read it carefully.

Has GPTZero itself tested WriteHuman?

Yes. On 30 July 2026 GPTZero published a hands-on review running one 800-word ChatGPT sample through WriteHuman and then through four detectors. ZeroGPT and Copyleaks both returned 0% AI, Originality.ai returned 63% confident human — and GPTZero returned 84% AI paraphrasing, with the verdict “We are highly confident this text was originally AI, but rewritten by AI or human.” Its conclusion was that WriteHuman “couldn’t bypass GPTZero as it claims.” That is one sample published by the detector being tested, so it is n=1 with an obvious interest.

Do those two sources contradict each other?

No, and that is the interesting part. GPTZero ran its test on 30 July 2026. WriteHuman’s own data for the July cycle records its worst GPTZero result of the four months at 60.8%. Two sources with opposite commercial incentives independently place late July as a weak period for WriteHuman against GPTZero. When adversaries agree, the finding is worth more than either claim alone.

Which detector is actually hardest for it?

It changes, which is itself the finding. In September 2026 WriteHuman’s data puts GPTZero at 71.8% — the lowest of its five detectors, below Winston at 84.1%, ZeroGPT at 77.9%, Copyleaks at 86.7% and Originality.ai at 91.4%. In June and July, Originality.ai was the hardest, at 46.3% and 53.6%. The commonplace that GPTZero is always the soft one does not hold for this tool in this data.

Does the type of writing matter?

More than the tool choice, on this evidence. In the September cycle WriteHuman’s bypass rate by category runs from 99.97% on academic essays and 100% on discussion-board posts down to 52.9% on business emails and 75.9% on application essays. The same product, the same month, the same detectors — and roughly a coin flip on one genre versus near-certainty on another.

Does GPTZero name WriteHuman anywhere else?

Yes. GPTZero’s February 2026 paper includes an appendix table naming nine bypass services, and WriteHuman is rated “Medium” there, alongside Grubby AI and TwainGPT. Read that column carefully: it is headed “bypassing ability on naive AI detection methods”, so it scores those tools against weak detectors, not against GPTZero itself.

About the author. Vlad Ivanov publishes Detection Drama, a site dedicated to AI-detection and text-humanization testing, with 245 published teardowns of detectors and humanizer tools including Turnitin, Pangram, GPTZero and WriteHuman. Profile: LinkedIn. This page reviews published evidence and reports no first-party test; corrections with a source are welcome.