rel=nofollow citations. The promoted call-to-action above is a separate StealthWriter affiliate placement with its own disclosure. Our companion page on Turnitin reached an unfavourable conclusion about the same product.Does Undetectable AI bypass GPTZero? Two detector companies have published figures pointing that way — and both of them sell against it.
The only tool that bypassed Pangram and Turnitin in 2026.
Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.
Try StealthWriter →Key Takeaways
- This is the one cell in this series where the evidence leans toward the tool — and it comes from two companies that sell against it.
- GPTZero’s own February 2026 paper rates Undetectable’s bypassing ability “High” — the only one of nine services to get that rating.
- Pangram’s August 2025 benchmark caught it 90.3% of the time — the lowest detection rate of all 20 tools in that table.
- Both are adverse-interest evidence: detector companies conceding a rival humanizer is comparatively hard to catch.
- Read the fine print, though. GPTZero’s “High” rating is explicitly against “naive AI detectors”, and its caption claims all nine are “ineffective against GPTZero.”
- And 90.3% is still nine times in ten caught. “Best of twenty” is not the same as “works”.
- Pangram’s DAMAGE paper separately rates Undetectable’s output quality L3, the lowest tier.
The only tool that bypassed Pangram and Turnitin in 2026.
Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.
Try StealthWriter →Does Undetectable AI bypass GPTZero?
This page is the exception in this series, and it is worth being upfront about why.
Across the Pangram and Turnitin clusters, the recurring finding was an absence: vendors claiming detector bypass with no method, no screenshot and no independent test. Here there is evidence — not from the vendor, but from two companies with every commercial reason to say the opposite.
That is adverse-interest evidence, and it is the most credible kind available in this market. It is also considerably weaker than the headline makes it sound, for two specific reasons we get to below.
This page reports no first-party test; everything below was fetched on 3 September 2026.
What GPTZero actually published
Appendix M of GPTZero’s own February 2026 paper contains something Turnitin has never published: a table naming humanizer tools individually.
Nine services, each with a rating. Undetectable is the only one marked “High”. Quillbot, StealthGPT, StealthWriter, HIX and GPTinf are all “Low”; Grubby AI, TwainGPT and WriteHuman are “Medium”.
The column header reads “bypassing ability on naive AI detection methods”. It is a rating of those tools against weak detectors — not against GPTZero.
And the caption states GPTZero’s own position outright: “These methods are all ineffective against GPTZero due to the four-tiered red teaming approach.”
So GPTZero is saying two things at once: Undetectable is the strongest of the nine against ordinary detectors, and none of the nine works against GPTZero specifically. Anyone quoting the “High” rating as proof it beats GPTZero is quoting it backwards.
Note also what that paper is: a GPTZero preprint, authored by GPTZero employees, not peer-reviewed. GPTZero has an obvious interest in the conclusion that everything fails against it.
What Pangram published
Pangram’s August 2025 benchmark scored twenty humanizers by how often Pangram caught their output. Most sit at 99–100%.
| Tool | Caught by Pangram |
|---|---|
| Undetectable AI | 90.3% — lowest of 20 |
| TwainGPT | 92.7% |
| Just Done | 93.5% |
| humanizeai.io | 93.8% |
| StealthGPT | 95.6% |
| Quillbot, Grammarly, Ahrefs and 8 others | 99–100% |
Bottom of a table of twenty is a real ranking, and it is consistent with GPTZero’s independent assessment. Two detector companies, different methods, same relative conclusion.
But 90.3% means caught nine times out of ten. Being the hardest of twenty to catch is not the same as not being caught. If a marker runs Pangram, the published expectation is still a flag.
Pangram’s DAMAGE paper separately rates Undetectable AI’s output quality L3, its lowest tier. Two caveats belong with that: the paper says it classifies on faithfulness and fluency “not on their effectiveness at bypassing AI detectors”, and it observes that less fluent output can be harder to detect. The L3 rating and the strong evasion ranking may be one fact seen from two directions — better at evading, worse to read.
Why you do not have to take anyone’s word for it
This is the structural difference between the GPTZero question and the Turnitin one, and it changes what you should demand of any vendor.
Turnitin sells no individual licences and shows students no AI report, so nobody outside an institution can produce a genuine screenshot. GPTZero has no such barrier. Anyone can paste up to 10,000 characters into gptzero.me with no account, no card and no signup. A free account raises that to 150,000 characters per scan with a 10,000-word monthly quota.
So when a humanizer claims to beat GPTZero and shows you nothing, that is a choice, not a constraint.
It also means you can check your own text. A genuine current GPTZero result shows a three-class classification — human, AI or mixed — a confidence category rather than a bare percentage, a probability breakdown across all three classes, and sentence-by-sentence highlighting. A screenshot showing only “X% AI”, or one displaying perplexity and burstiness, is either fabricated or dates to roughly 2023.
How reliable is GPTZero itself?
Worth knowing, because it bounds what any of this proves.
GPTZero’s own FAQ concedes: “Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation.” Its developers page simultaneously claims a “Paraphraser Shield” means “even if AI content has been altered to look more human-like, GPTZero can detect it.” Both are live.
The two detector vendors publish opposite numbers on the same question. GPTZero’s paper reports 93.5% recall on 1,000 texts run through nine bypasser services, and puts Pangram at 49.7%. Pangram’s DAMAGE paper reports GPTZero at 60.04% on humanized text, and 34.53% at GPTZero’s own default threshold. Each benchmark shows its publisher winning.
On false positives, GPTZero was one of seven detectors in Liang et al. (2023), which reported a 61.22% average false positive rate across those seven on 91 TOEFL essays. No per-detector figure was ever published, so that number is not GPTZero’s. GPTZero now claims it has cut its own TOEFL false positive rate to 1.1% — a figure no third party has verified.
What should you take from this?
Of the twenty tools in this series, Undetectable AI has the best evidence against GPTZero, and it is not close. Two rival detector companies independently rank it the hardest of the field to catch, which is the only kind of favourable evidence in this market worth anything.
It is still not a guarantee, and the gap between those two statements is where people get into trouble. GPTZero’s own paper says all nine named tools fail against it. Pangram still caught nine samples in ten. The quality rating on the output is the lowest tier. And there is no published screenshot from the vendor of a detector that is free to test.
The honest position: better odds than its rivals, on a question you can settle yourself for nothing in two minutes, against a detector that is unreliable in both directions.
If your real worry is being wrongly flagged on your own writing, keep your drafting history — it costs nothing and it survives a detector being wrong. It matters most for the writers detectors treat worst: ESL writers and AI detection and false positives for neurodivergent students. Check before you submit with the best pre-submission check and the Turnitin self-check, and strip the obvious tells by hand first via what to remove before using an AI humanizer.
For the same product against other detectors, see Turnitin and Pangram, and the Pangram review. Also in this cluster: SuperHumanizer.
Get the free prompt pack instead
The manual rewriting prompts that lower AI signals without a subscription — and without betting a submission on a number nobody has reproduced. Free, no card.
Frequently asked questions
Does Undetectable AI bypass GPTZero?
It is the strongest candidate in this market, and that assessment comes from its competitors rather than its marketing. GPTZero’s own paper rates its bypassing ability “High” — alone among nine named services — and Pangram’s benchmark caught it less often than any of the twenty tools tested. Neither is a clean pass, and both carry caveats that matter.
Why does competitor evidence count for more here?
Because of the direction it runs. A vendor saying its own tool works is marketing. A detector company publishing that a rival humanizer is comparatively hard to catch is conceding something against its own commercial interest. Two separate detector companies have now done that about Undetectable AI, independently.
What exactly does GPTZero’s rating mean?
Less than it first appears. Appendix M of GPTZero’s February 2026 paper lists nine bypass services with a “bypassing ability” column, and Undetectable is the only one marked High. But the column header reads “on naive AI detection methods” — it rates those tools against weak detectors, not against GPTZero. The table’s own caption says: “These methods are all ineffective against GPTZero due to the four-tiered red teaming approach.”
And the Pangram figure?
Pangram’s August 2025 benchmark reports catching Undetectable AI output 90.3% of the time. That is the lowest figure in a table of twenty tools, most of which sit at 99–100%. It is meaningful as a relative ranking. It is not a pass: nine times in ten, Pangram still caught it.
Is there any first-party evidence from Undetectable AI itself?
Nothing verifiable of the kind that matters. And unlike Turnitin, there is no excuse for that — GPTZero accepts 10,000 characters with no account and no payment, so any vendor could publish a genuine screenshot in minutes.
What does GPTZero say about paraphrased text generally?
Two contradictory things, both currently published. Its FAQ concedes: “Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation.” Its developers page claims a “Paraphraser Shield” means “even if AI content has been altered to look more human-like, GPTZero can detect it.”
What about output quality?
Pangram’s DAMAGE paper rates Undetectable AI “L3”, the lowest of three quality tiers. Two caveats: the paper explicitly rates faithfulness and fluency “not… effectiveness at bypassing AI detectors”, and it separately notes that less fluent output can be harder to detect. So the L3 rating and the strong evasion ranking may be the same fact seen from two sides.
So should I use it?
That is not what this page decides. What it establishes is that the evidence for this specific tool against this specific detector is better than for any of its rivals, that it still falls well short of a guarantee, and that you can verify your own output free in about two minutes rather than trusting anyone’s number.
