Does StealthGPT Bypass Pangram? What the Published Data Actually Shows

Published:

Updated:

Vlad Ivanov
Vlad Ivanov · LinkedIn
Runs detectiondrama.com, where he has tested and documented 40+ AI humanizers and detectors against Turnitin, GPTZero, Originality.ai and Pangram since 2024. This piece is an evidence review of published third-party data, not a first-party test.
Last updated: July 2026

Does StealthGPT bypass Pangram? No — not reliably. Pangram’s own published humanizer benchmark scores StealthGPT at 95.6% detected, and an independent University of Chicago study names StealthGPT specifically as a tool Pangram holds up against.

Detection Drama · Free Download

Want to bypass Turnitin in 2026? Grab the free prompt pack.

Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.

Send me the free prompts →
Free · No credit card · Straight to your inbox

Key Takeaways

  • Pangram catches 95.6% of StealthGPT output in its own August 2025 benchmark of 20 humanizers — StealthGPT ranked 15th out of 20 for evasion.
  • Undetectable AI is the weakest humanizer against Pangram at 90.3% detected — still nowhere near a bypass.
  • An independent University of Chicago study (Jabarian & Imas, 1,992 human + 1,992 AI texts) found Pangram “robust to humanizer tools (e.g., StealthGPT)” — naming StealthGPT by name.
  • Pangram’s claimed false positive rate is 1 in 10,000, roughly 38x lower than a typical commercial detector.
  • Pangram CTO Bradley Emi says the ~4% that escapes is mostly text “so badly garbled it barely resembles English” — output a human grader spots instantly.
  • The best available numbers are vendor-published and 11 months old. Treat the 95.6% as directional, not gospel.

Methodology note: This is an evidence review, not a first-party test. I did not run StealthGPT output through Pangram myself. Every number below is published by a named source and linked to its primary document, with the funding and recency of each source stated plainly so you can weight it yourself. Where the data is a vendor grading its own product, I say so.

Detection Drama · Free Download

Want to bypass Turnitin in 2026? Grab the free prompt pack.

Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.

Send me the free prompts →
Free · No credit card · Straight to your inbox

Has anyone actually tested StealthGPT against Pangram?

Yes — but not the people you would expect, and not where anyone thought to look. Search “StealthGPT vs Pangram” and you get nothing useful: humanizer review sites test StealthGPT against Turnitin and GPTZero, and detector review sites test Pangram against other detectors. The two lanes never cross.

The test exists anyway. It is sitting in Pangram’s own product blog, published August 27, 2025 by co-founder and CTO Bradley Emi, and it names StealthGPT in a table of 20 humanizers. It has been public for nearly a year and the humanizer-review industry has almost entirely ignored it — which is unsurprising, given what it says.

What does Pangram’s benchmark say about StealthGPT?

Pangram reports 95.6% accuracy on StealthGPT output. In plain terms: feed 100 StealthGPT-humanized passages into Pangram and roughly 96 come back flagged as AI.

That places StealthGPT in the lower-middle of the pack — better at evading than Grammarly or Quillbot (both 100% detected, i.e. useless for evasion), but worse than Undetectable AI, which at 90.3% is the single weakest performer against Pangram in the whole table. Here is the full published set, ranked by how often Pangram caught them:

HumanizerPangram accuracyRead as
Undetectable AI90.3%Best evader tested — still caught ~9 times in 10
TwainGPT92.7%Caught ~9 times in 10
Just Done93.5%Caught ~9 times in 10
humanizeai.io93.8%Caught ~9 times in 10
StealthGPT95.6%Caught ~24 times in 25
DIPPER97.6%Caught ~49 times in 50
Writesonic AI98.1%Near-total detection
Scribbr99.0%Near-total detection
GPTinf99.2%Near-total detection
Bypass GPT99.7%Near-total detection
Ahrefs, aihumanizer.com, Ghost AI, Grammarly, humanizeai.pro, Quillbot, Semihuman AI, Smodin, Surfer SEO, surgegraph.io100.0%Zero evasion

Source: Pangram Labs, “How well does Pangram perform on humanizers?”, updated August 27, 2025. Sample size and per-tool methodology are not disclosed in the post.

Weight this correctly. That table is Pangram grading its own homework. A detector vendor publishing a benchmark showing its detector wins has an obvious commercial interest in the result, and the post does not disclose sample sizes, passage lengths, prompt sources, or which StealthGPT mode was used. It is the best available data on this exact question — and it is still marketing collateral. The independent corroboration below is what actually makes it credible.

Does independent research back that up?

Partly, and importantly — one independent study names StealthGPT directly.

Economists Brian Jabarian and Alex Imas at the University of Chicago’s Becker Friedman Institute compared four detectors — Pangram, GPTZero, Originality.ai, and open-source RoBERTa — across 1,992 human texts written pre-2020 and 1,992 AI-generated texts spanning blogs, reviews, resumes, news, and novels. Their working paper states:

“Pangram’s performance largely holds up on very short passages (< 50 words) and is robust to ‘humanizer’ tools (e.g., StealthGPT), the performance of other detectors becomes case-dependent.”

That is an independent academic source, using StealthGPT as its named example of a humanizer that does not defeat Pangram. The same study found Pangram “dominates the other detectors across all thresholds” and was the only one meeting a strict false-positive cap (FPR ≤ 0.005) without losing detection power.

A separate University of Maryland benchmark (Russell et al.) tested detectors on humanized text and put Pangram at 97% accuracy against GPTZero at 46%, Fast-DetectGPT at 23%, and Binoculars at 7%. And Jabarian and Imas’s follow-up work found Pangram was the only one of four commercial detectors whose performance stayed robust against humanizers at all.

StealthGPT vs Pangram key numbers: 95.6% of StealthGPT output caught by Pangram, 90.3% Undetectable AI, 97% Pangram accuracy on humanized text, 46% GPTZero, 1 in 10,000 false positive rate, 20 humanizers benchmarked
The published numbers on StealthGPT vs Pangram. Sources: Pangram Labs humanizer benchmark (Aug 2025); Russell et al., University of Maryland.
95.6%of StealthGPT output caught by Pangram (Pangram, Aug 2025)
90.3%Undetectable AI — the weakest humanizer tested
97%Pangram accuracy on humanized text vs GPTZero’s 46% (Russell et al., UMD)
1 in 10,000Pangram’s claimed false positive rate
3,984texts compared in the UChicago detector study
$0.0228Pangram cost per correctly flagged passage vs $0.0575 for GPTZero

Why is Pangram so much harder to bypass than Turnitin or GPTZero?

Because most humanizers are built to defeat a different kind of detector than Pangram is.

Classic detectors lean on perplexity — how statistically predictable each next word is. Humanizers game that directly: swap synonyms, vary sentence length, inject odd punctuation, and the perplexity score moves. That is why bypass rates against older detectors can look impressive in vendor marketing.

Pangram trains against humanizer output as a data-augmentation strategy rather than measuring predictability alone. Its classifier, EditLens, was published in a peer-reviewed ICLR 2026 paper, and the company’s DAMAGE paper audited 19 humanizer and paraphraser tools specifically to learn their fingerprints. Pangram even reports training an adversarial model tuned against its own predictions and finding its cross-humanizer generalization survived the attack.

The result is a counterintuitive finding that Emi states outright: the better a humanizer is at sounding fluent, the more reliably Pangram catches it. Fluent output stays inside the statistical neighbourhood Pangram was trained on. It is the incoherent output that slips through.

What about the 4.4% that gets through?

This is the part that decides the practical answer, and it is the part every “95.6% means a 4.4% success rate!” reading gets wrong.

Pangram explains why its own number is not 100%, and the explanation is unflattering for humanizer users. Emi gives two reasons. First, Pangram has internal models that detect humanizers at near-perfect accuracy but ships the weaker one on purpose, because the stronger models raise the false positive rate and the company would rather miss AI text than flag a real student. Second — and this is the one that matters:

“Most of the cases in which Pangram does not catch humanized output, the text is so badly garbled and obfuscated that it barely resembles English. These cases are easy to spot by eye.”

So the 4.4% is not a hidden setting or a clever prompt. It is the tail of the distribution where the output stopped being readable. You can beat the machine by submitting something a human marker will flag in one paragraph — which is not beating anything. This is the same trap as humanizers that mangle facts and humanizers that wreck a thesis statement: the evasion and the quality move in opposite directions.

Pangram also publishes the visible tells it looks for, which double as a checklist for spotting humanized text by eye: tortured phrases (“counterfeit consciousness” for “artificial intelligence”), unnatural spacing, repeated synonym substitutions, and non-standard Unicode characters such as U+2009 thin spaces used to confuse tokenizers.

How current is this data?

Not very, and that is the honest limit on everything above.

The benchmark is dated August 27, 2025 — roughly eleven months old as of this writing. Both sides have shipped since. Pangram 3.3 landed in May 2026, after Pangram reportedly struggled with GPT-5.4 output in March 2026 and then recovered. StealthGPT has not stood still either. The 95.6% is a snapshot of a moving fight.

Which way did it move? Both directions are plausible. Detector-side, Pangram has kept publishing research and shipping models. Humanizer-side, StealthGPT has had eleven months to tune against exactly this benchmark. Nobody has published a fresh number — and until someone does, anyone quoting a 2026 StealthGPT-vs-Pangram figure is either extrapolating or making it up.

EvidenceSourceIndependent?Date
StealthGPT 95.6% detectedPangram humanizer benchmarkNo — vendor self-gradedAug 2025
Pangram “robust to humanizers (e.g., StealthGPT)”UChicago BFI (Jabarian & Imas)Yes — academicAug 2025
Pangram 97% vs GPTZero 46% on humanized textUMD (Russell et al.)Yes — academicJan 2025
StealthGPT fails Turnitin (86% AI), Originality.ai (100% AI)Phrasly reviewNo — competitor humanizer2026
Any 2026 StealthGPT-vs-Pangram testDoes not exist

So what is the honest verdict?

Verdict: On the best available evidence, StealthGPT does not bypass Pangram. Pangram’s own benchmark catches it 95.6% of the time; the one independent study that names StealthGPT uses it as the example of a humanizer Pangram withstands. The data is vendor-flavoured and eleven months old, so treat it as directional — but there is no published evidence pointing the other way, and StealthGPT’s record against weaker detectors (86% AI on Turnitin, 100% AI on Originality.ai) gives no reason to expect it does better against the hardest one.

If your institution runs Pangram, no humanizer in the published table gets you under a coin flip — the best of them, Undetectable AI, still gets caught nine times in ten. The tool choice is not the variable. The strategy is.

That points somewhere more boring and more effective: lowering your AI score without humanizer tricks, keeping drafting evidence that proves authorship, and understanding how often detectors are wrong in the other direction — because a documented writing process survives a false positive, and a humanizer does not.

Frequently asked questions

Does StealthGPT bypass Pangram in 2026?

There is no published 2026 test. The most recent data is Pangram’s August 2025 benchmark showing 95.6% of StealthGPT output detected. Pangram has shipped version 3.3 since (May 2026) and StealthGPT has also updated, so the current number is unknown — but no published evidence suggests StealthGPT has gained the upper hand.

Which AI humanizer is hardest for Pangram to detect?

Of the 20 tools in Pangram’s published benchmark, Undetectable AI performed best at 90.3% detected — meaning Pangram still caught roughly nine of every ten passages. TwainGPT (92.7%) and Just Done (93.5%) followed. No tested humanizer got below 90%.

Is Pangram’s benchmark trustworthy if Pangram published it?

Treat it as directional rather than definitive. It is vendor-published, self-graded, and omits sample sizes and methodology. It gains credibility from independent corroboration: the University of Chicago study reached the same directional conclusion and named StealthGPT specifically, and a University of Maryland benchmark independently put Pangram at 97% on humanized text.

Why does Pangram catch humanizers when GPTZero and Turnitin do not?

Most detectors measure perplexity — how predictable each word is — which humanizers are explicitly built to disrupt. Pangram trains directly on humanizer output using data augmentation, so evasion patterns become a signal rather than noise. In UMD testing, Pangram hit 97% on humanized text where GPTZero managed 46%.

What is the 4.4% of StealthGPT text that Pangram misses?

According to Pangram CTO Bradley Emi, most missed cases are text “so badly garbled and obfuscated that it barely resembles English” — easy for a human to spot even though it is hard to catch algorithmically. The output that evades the detector is generally output that fails a human read.

Does Pangram flag human writing as AI?

Pangram claims a false positive rate of 1 in 10,000, roughly 38 times lower than typical commercial detectors. The University of Chicago study found it was the only detector of four to meet a strict FPR cap of 0.005 without losing detection power. That is a strong claim relative to competitors, but no detector is infallible.

Can Pangram tell that text was run through a humanizer at all?

Yes — Pangram publishes the tells it looks for, including tortured phrases (odd synonym swaps like “counterfeit consciousness” for “artificial intelligence”), unnatural spacing, repeated synonym substitutions, and non-standard Unicode characters such as U+2009 thin spaces inserted to confuse tokenizers.

Should I use a different humanizer against Pangram?

The published data does not support any tool as a reliable Pangram bypass — every one of the 20 humanizers tested was caught at least 90% of the time. Tool choice is not the deciding variable when the whole category performs this way against this detector.