rel=nofollow citations. The promoted call-to-action above is a separate StealthWriter affiliate placement carrying its own disclosure.Does StealthGPT bypass GPTZero? About half the time — which is both far more than most tools in this series can show, and far less than the 0-out-of-30 it advertises.
The only tool that bypassed Pangram and Turnitin in 2026.
Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.
Try StealthWriter →Key Takeaways
- This is the one cell in this series where a rigorous independent academic study directly answers the question — and the answer is a qualified yes.
- The UChicago BFI paper ran 1,992 matched passages through StealthGPT’s own rewrite endpoint. GPTZero’s false negative rate went from 0–7% on raw AI text to 35–77% on StealthGPT output — it misses roughly half.
- Pangram, on the identical text, stayed at 0–5%. StealthGPT beats the detector every dataset calls the easiest and loses to the one every dataset calls the hardest.
- StealthGPT advertises “GPTZero: 95% human average (0/30 detections)” — undated, no method, no source text.
- Its own published 2026 screenshot returns 75% human, not 95% — with GPTZero stating it is only “moderately confident” and putting the text at AI 25%. We verified the screenshot forensically; it is genuine.
- Every 2026 screenshot on that page scans the same single coffee-brewing passage — one text, one run per detector, no repetition.
- Between 20 August and 3 September 2026 its pricing switched from weekly to monthly and the money-back-if-detected guarantee vanished from the pricing page.
The only tool that bypassed Pangram and Turnitin in 2026.
Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.
Try StealthWriter →Does StealthGPT bypass GPTZero?
Forty-three articles into this series, the answer to “does this tool bypass that detector” has almost always been the same: nobody knows, because nobody independent has tested it.
This cell is different. StealthGPT is the humanizer used in the most rigorous published study of humanized-text detection in existence, and that study evaluated GPTZero directly. For once there is a real answer.
The answer is partly — and the gap between that and what the company sells is the story.
This page reports no first-party test of our own. Every source below is dated and attributed, and everything vendor-side was fetched on 3 September 2026.
What does the strongest study actually find?
The University of Chicago BFI working paper by Brian Jabarian and Alex Imas (August 2025) built a corpus of 1,992 human passages matched with AI-generated equivalents across six genres — news, blogs, consumer reviews, novels, restaurant reviews and résumés — generated by four frontier models. Then, in the authors’ words: “We fed every AI passage through the StealthGPT default rewrite endpoint and re-scored the outputs.”
What follows is a false negative rate: the share of AI text a detector fails to catch. Higher is worse for the detector, better for the humanizer.
| Detector | Raw AI text (FNR) | After StealthGPT (FNR) | Read |
|---|---|---|---|
| GPTZero | 0.00–0.07 | 0.35–0.77 | Misses roughly half |
| Originality.ai | low | 0.02–0.21 | Degraded, still working |
| Pangram | low | 0.00–0.05 | Essentially unaffected |
The paper’s own sentence is blunt: “GPTZero largely loses its capacity to detect AI-generated text, with FNR scores around 0.50 and above across most genres and LLM models” — while Pangram is described as “largely robust to humanizer.”
Can you check this yourself?
Turnitin sells no individual licences and shows students no AI report, so nobody outside an institution can produce a genuine Turnitin screenshot. GPTZero removes that excuse entirely. Anyone can paste up to 10,000 characters into gptzero.me right now with no account, no card and no signup; a free account raises the per-scan limit to 150,000 characters with a 10,000-word monthly quota, and paid tiers start at $9.99 a month.
So when a humanizer claims to beat GPTZero and shows you nothing, that is a choice, not a constraint.
It also means you can check your own text. A genuine current GPTZero result shows a three-class classification — human, AI or mixed — a confidence category rather than a bare percentage, a probability breakdown across all three classes, and sentence-by-sentence highlighting. An image showing only “X% AI”, or one displaying perplexity and burstiness, is either fabricated or dates to roughly 2023.
What does StealthGPT claim, and does its own screenshot hold up?
The headline number appears on the homepage and again on its explainer post: “GPTZero: 95% human average (0/30 detections).” Thirty tests is a real sample size and StealthGPT deserves more credit than the many vendors in this series who publish no number at all. But no date, no source texts, no model, no scan mode and no screenshots are attached to it anywhere we could find.
Then, further down that same page, the company runs what it calls a bonus round:
“As a bonus round we decided to run a test against the best of the best, the 2026 model in GPTZero. Being the toughest detector out there, StealthGPT passed with flying colors at solid 75%. See the screenshot below.”stealthgpt.ai, read 3 September 2026
Two things follow. First, StealthGPT calls GPTZero “the toughest detector out there” — a description no dataset supports, and one its own results contradict. Second, the number it shows is 75%, twenty points below the 95% average printed higher up the same page, and it is the only one of the two with evidence attached.
Is the screenshot real?
Yes, and that is worth saying plainly, because in this series it is unusual. We downloaded the image and inspected it rather than taking the caption’s word for it.
A fabricated or recycled GPTZero screenshot fails specific tests. This one passes them. It shows a genuine three-class breakdown — AI 25% / Mixed 0% / Human 75%, a confidence category rather than a bare percentage, a model version stamp reading Model 4.4b, the current Basic Scan layout, and a signed-out Sign up / Log in state consistent with GPTZero’s free public tier. Nothing about it is stale or invented.
There is one more thing the screenshots show collectively. Every 2026 detector scan on that page — Originality.ai, Conch, Content at Scale, Copyleaks, Winston and GPTZero — is run on the same single passage, a generic coffee-brewing how-to of roughly 300 words. One source text. One run per detector. No repetition, no genre variation, no timestamps. Set that against a study whose results move 42 points on genre alone.
Has GPTZero itself said anything about StealthGPT?
It has, and this is the second piece of adverse-interest evidence. GPTZero’s own February 2026 paper contains something Turnitin has never published: an appendix table naming nine bypass services individually. StealthGPT is one of them, rated “Low”, alongside Quillbot, StealthWriter, HIX and GPTinf. Grubby AI, TwainGPT and WriteHuman are “Medium”; only Undetectable is “High”.
So GPTZero asserts StealthGPT is ineffective against it, and the BFI paper measures GPTZero missing between a third and three-quarters of StealthGPT output. Those cannot both be right. One is a vendor’s claim about its own product; the other is a third-party measurement with a published corpus and a stated method. We weight them accordingly.
How reliable are the numbers on either side?
GPTZero’s own FAQ concedes: “Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation.” Its developers page simultaneously claims a “Paraphraser Shield” means “even if AI content has been altered to look more human-like, GPTZero can detect it.” Both statements are currently published.
The two leading detector vendors publish opposite numbers on the same question. GPTZero’s February 2026 paper reports 93.5% recall on 1,000 texts run through nine bypasser services, and puts Pangram at 49.7%. Pangram’s DAMAGE paper reports GPTZero at 60.04% on humanized text, and 34.53% at GPTZero’s own default threshold. Each benchmark shows its publisher winning, and neither is peer-reviewed.
On false positives, GPTZero was one of seven detectors in Liang et al. (2023), which reported a 61.22% average false positive rate across those seven on 91 TOEFL essays. No per-detector figure was ever published, so that number is not GPTZero’s. GPTZero now claims it has cut its own TOEFL false positive rate to 1.1% — a figure no third party has verified.
What happens against the harder detectors?
This is where the picture completes, and it is the part the marketing omits.
Pangram’s August 2025 benchmark scored twenty humanizers by how often Pangram caught their output. StealthGPT sat at 95.6% caught — fifth-easiest of twenty, well behind Undetectable AI’s 90.3%. Pangram is a commercial detector vendor scoring tools it competes with, and the post discloses no sample size or false-positive rate, so treat it as directional. But the independent BFI paper reports the same shape from a different method: Pangram’s miss rate on StealthGPT text is 0.00–0.05.
And on Turnitin, the one independent test that exists went the other way entirely: when Turnitin shipped bypasser detection in August 2025, a StealthGPT document that had scored 0% likely-AI was re-scored at 72% — the largest swing of any tool in that run. We cover it in full in does StealthGPT bypass Turnitin.
The defensible summary across all three: StealthGPT beats the detector every dataset agrees is the softest, and loses to the two that are not. “99% AI-detection bypass rate” is not a description of that.
What changed on the pricing page?
Something worth recording, because it happened inside two weeks.
On 20 August 2026, when we fetched the site for the Turnitin article, StealthGPT sold weekly: $9.99 a week for 50 requests and $19.99 for 100, both promoted at $1 for a first week, plus a 99-cent student offer, and a homepage FAQ stating that subscriptions came with a money-back guarantee if you were detected.
On 3 September 2026, the pricing page lists Pro at $29.99 a month (50,000 words, 1,000 per request), Max at $49.99 (200,000 words, 2,000 per request) and Enterprise at $179.99 (1,000,000+ words), with a yearly option advertised as saving up to 17.5%. There is no free tier. And the words guarantee, refund and money back do not appear on that page at all — only “Cancel anytime.”
We are not claiming the guarantee was withdrawn everywhere; we are recording that it is not on the page where you now pay.
What should you take from this?
StealthGPT does something real to GPTZero. That is not a marketing claim here; it is the finding of the most careful study anyone has published, using StealthGPT’s own endpoint, and it survives the fact that the study had no reason to flatter the company.
What it does not do is what the checkout says. 0 out of 30 is not the same as missing half. The company’s own screenshot returns 75% human at moderate confidence, not 95%. Its own test suite is one coffee-brewing paragraph. And the detector it beats is the one that free public access makes easiest to verify and every dataset agrees is easiest to beat — while Pangram catches it 95.6% of the time and Turnitin re-scored it at 72% likely-AI once it went looking.
If GPTZero is genuinely the only detector standing between you and a problem, this tool moves the odds. If it is Turnitin, Pangram or Originality.ai, the evidence says it does not.
If your real worry is being wrongly flagged on your own writing, keep your drafting history — it costs nothing and it survives a detector being wrong. It matters most for the writers detectors treat worst: ESL writers and AI detection and false positives for neurodivergent students. Check before you submit with the best pre-submission check and the Turnitin self-check, and strip the obvious tells by hand first via what to remove before using an AI humanizer.
For the product itself see our StealthGPT review, and the same question against Turnitin and Pangram. Also in this cluster: GPTHuman, Undetectable AI, SuperHumanizer.
Get the free prompt pack instead
The manual rewriting prompts that lower AI signals without a subscription — and without betting a submission on a number nobody has reproduced. Free, no card.
Frequently asked questions
Does StealthGPT bypass GPTZero?
Roughly half the time, on the strongest evidence available. The University of Chicago BFI working paper by Jabarian and Imas fed 1,992 AI passages through StealthGPT’s default rewrite endpoint and re-scored them. GPTZero’s false negative rate rose from 0.00–0.07 on unhumanized AI text to 0.35–0.77 on StealthGPT output, depending on genre and model. The paper’s own summary is that GPTZero “largely loses its capacity to detect AI-generated text.” That is a real effect, and it is nowhere near the 0/30 the vendor advertises.
Why is this article more confident than the others in this series?
Because for once the evidence exists. Most humanizers have never been tested by anyone independent. StealthGPT was used as the humanizer in a serious academic study with a published corpus, a stated endpoint and per-genre results. We are not relying on the vendor’s numbers, and we are not guessing.
What does StealthGPT itself claim about GPTZero?
Two different things on the same page. A homepage and blog banner claims “GPTZero: 95% human average (0/30 detections)” with no date, no source text and no methodology. Further down the same blog post, a 2026 bonus test against GPTZero returns 75% human, described as passing “with flying colors.” Those two numbers are twenty points apart and the lower one is the one with a screenshot attached.
Is the StealthGPT screenshot genuine?
Yes. We downloaded and inspected it. It shows GPTZero’s current three-class result UI — AI 25% / Mixed 0% / Human 75% — a model version stamp reading Model 4.4b, a Basic Scan header, and the signed-out Sign up / Log in state consistent with GPTZero’s free public tier. Nothing about it is fabricated or recycled. What it does not support is the claim printed above it.
So what is wrong with the test?
Its design. Every 2026 detector screenshot on that page scans the same single passage — a generic coffee-brewing how-to of roughly 300 words. One source text, one run per detector, no repetition, no genre variation, no date stamp. The BFI paper shows exactly why that matters: GPTZero’s miss rate on StealthGPT output swings from 35% to 77% depending on genre alone.
Does GPTZero name StealthGPT?
Yes. GPTZero’s February 2026 paper includes an appendix table naming nine bypass services, and StealthGPT is one of them, rated “Low”. That rating is widely misread. The column header says “bypassing ability on naive AI detection methods” — it scores those tools against weak detectors, not against GPTZero. The caption then asserts that all nine are ineffective against GPTZero, which the BFI data contradicts. Both documents are preprints, and each publisher’s own detector wins in its own paper.
What does Pangram find?
Pangram’s August 2025 benchmark of twenty humanizers caught StealthGPT output 95.6% of the time, placing it fifth-easiest to catch of twenty. The BFI paper agrees from the other direction, reporting Pangram’s false negative rate on StealthGPT-humanized text at 0.00–0.05. Both sources are consistent: whatever StealthGPT does to GPTZero, it does not do to Pangram.
What does it cost, and is there still a guarantee?
As of 3 September 2026 the pricing page lists Pro at $29.99 a month for 50,000 words, Max at $49.99 for 200,000 and Enterprise at $179.99 for a million or more, with a yearly option advertised as saving up to 17.5%. There is no free tier. This is a change: on 20 August 2026 the same product was sold weekly at $9.99 and $19.99 with a $1 first week and a 99-cent student offer. The words guarantee, refund and money back do not appear on the current pricing page at all.
