Does Phrasly Bypass GPTZero? Its Changelog Says Not Reliably (2026)

Published:

Updated:

Disclosure. Phrasly is not a Detection Drama affiliate and this page carries no revenue path to it; outbound links are rel=nofollow citations. The promoted call-to-action above is a StealthWriter affiliate placement carrying its own disclosure, and is unrelated to this review.

Does Phrasly bypass GPTZero? Its own engineering changelog records GPTZero catching it three separate times, and no pass rate at all.

Detection Drama · 2026 Verdict

The only tool that bypassed Pangram and Turnitin in 2026.

Most humanizers clear one detector and get caught by the other. StealthWriter is the one that gets past both — run Ghost 5.2 Pro at level 7–8, section by section, and re-check before you submit.

Try StealthWriter →
Free plan to start · No credit card · Paid from $20/mo
Or get the free prompt pack first →
Affiliate link — I earn a commission if you upgrade. I only point at tools I actually run.

Key Takeaways

  • Phrasly keeps a real versioned engineering changelog — fifteen dated entries back to August 2023. In this market that is close to unique, and it is the reason this page can be written from records rather than marketing.
  • Every GPTZero entry in it is a repair. Three separate dated entries log GPTZero’s detection catching Phrasly and Phrasly patching afterwards.
  • 21 Jan 2025: “Adapted the model to GPTZero’s paraphrasing-detection update.” 17 Jun 2025: regressions against “GPTZero’s updated paraphrasing detection” and “GPTZero’s AI-editing detection.” 9 Jan 2026: another paraphrasing-detection regression.
  • For Turnitin it published a dated, numbered compatibility review — 100 documents, all 0%. For GPTZero there is no compatibility review at all, and no number anywhere.
  • That is backwards. Turnitin results cannot be obtained by any vendor. GPTZero is free, instant and unlimited to test — so it is the one detector where a hundred-document run would actually be possible.
  • GPTZero tested Phrasly itself in December 2024 and flagged the output as “likely AI”, specifically labelled “AI paraphrased”, calling the results “inconsistent and misleading.” That is one sample, published by the detector.
  • Its homepage has since repositioned: output scores well not because it’s designed to trick them, but because quality writing naturally resembles human writing.”

By Vlad Ivanov — publisher of Detection Drama, tracking AI-detector and humanizer behaviour across 245 published tests and teardowns. Last updated: September 2026

Does Phrasly bypass GPTZero - 2026 evidence review

Does Phrasly bypass GPTZero?

For once this can be answered from records rather than claims, because Phrasly keeps something almost nobody else in this market does: a real versioned engineering changelog, fifteen dated entries running back to August 2023.

That changelog mentions GPTZero three times. Every one of them is a repair.

This page reports no first-party test. Everything below was fetched on 21 September 2026 and quoted verbatim.

What does the changelog actually record?

Every GPTZero entry in Phrasly's changelog is a repair

Three dated entries, each logging GPTZero’s detection catching Phrasly’s output and Phrasly adapting afterwards.

DateEntryShape
21 Jan 2025“Adapted the model to GPTZero’s paraphrasing-detection update.”Repair
17 Jun 2025Resolved regressions against “GPTZero’s updated paraphrasing detection” and “GPTZero’s AI-editing detection”Repair ×2
9 Jan 2026“Resolved a regression against GPTZero’s updated paraphrasing detection that affected some texts.”Repair
Any dateGPTZero compatibility review with a pass rateDoes not exist

Source: phrasly.ai/changelog, all fifteen entries read 21 September 2026. No sample size or numerical result accompanies any GPTZero entry.

None of these is a marketing document. A changelog entry saying you resolved a regression is written for engineering reasons, by people recording that something broke. That is what makes it worth more than a guarantee.

Read in sequence, the entries describe an arms race Phrasly is not winning outright. GPTZero updates its paraphrasing detection; Phrasly adapts. GPTZero ships AI-editing detection; Phrasly resolves a regression. GPTZero updates paraphrasing detection again; Phrasly resolves another. Three rounds across twelve months, with the detector moving first every time.

Why is the missing review the loudest part?

Because of which detector it is missing for.

Phrasly published a dated compatibility review for Turnitin on 27 August 2025, timed to Turnitin’s own product update and claiming 100 documents, all returning a 0% AI score. We examined that claim in does Phrasly bypass Turnitin. Its central problem is access: Turnitin sells no individual licences and its AI report is instructor-facing, so no vendor can lawfully produce a hundred such reports.

GPTZero is the exact opposite. Free, instant, no account, up to 10,000 characters a scan. It is the one major detector where a hundred-document run is genuinely achievable by anybody, including Phrasly.

So the company ran a large numbered study on the detector it could not possibly have tested at that scale, and published no study at all on the detector it could.

Can you check GPTZero yourself?

GPTZero is free to test - 10,000 characters with no account

Turnitin sells no individual licences and shows students no AI report, so nobody outside an institution can produce a genuine Turnitin screenshot. GPTZero removes that excuse entirely. Anyone can paste up to 10,000 characters into gptzero.me right now with no account, no card and no signup; a free account raises the per-scan limit to 150,000 characters with a 10,000-word monthly quota, and paid tiers start at $9.99 a month.

So when a humanizer claims to beat GPTZero and shows you nothing, that is a choice, not a constraint.

It also means you can check your own text. A genuine current GPTZero result shows a three-class classification — human, AI or mixed — a confidence category rather than a bare percentage, a probability breakdown across all three classes, and sentence-by-sentence highlighting. An image showing only “X% AI”, or one displaying perplexity and burstiness, is either fabricated or dates to roughly 2023.

What happened when GPTZero tested it?

On 23 December 2024, GPTZero published its own hands-on review of Phrasly. It generated a ChatGPT essay introduction about the Roman Empire, ran it through Phrasly’s aggressive humanizer setting, which produced five different rewrites, and scanned the results.

The original ChatGPT text scored 100% confidence as AI. After Phrasly, GPTZero flagged the rewritten paragraph as likely AI and specifically labelled it “AI paraphrased” — the Paraphraser Shield firing — with inconsistent results across the five drafts. Its conclusion was that “both the Phrasly AI Humanizer and Phrasly’s AI Detector are inconsistent and misleading when it comes to the outputs and promises of their product.”

Take that for exactly what it is. One source paragraph, five rewrites, published by the detector under test on a page arguing its own superiority. Alone it would prove little. What gives it weight is direction: it points the same way as Phrasly’s own changelog, which was written for engineers rather than for readers. A vendor’s internal record and a competitor’s test agreeing is the strongest signal available in this market.

Has Phrasly changed what it claims?

Yes, and the shift is worth recording accurately rather than mocking.

Its current homepage FAQ reads: “The result is better writing that also scores well on platforms like Turnitin and GPTZero, not because it’s designed to trick them, but because quality writing naturally resembles human writing.” It adds that users should treat Phrasly as “a refinement tool, not a substitute for their own thinking” and “follow your institution’s policies.”

Its Terms disclaim the outcome in capitals: “WE CANNOT GUARANTEE THAT THE TEXT GENERATED BY OUR SERVICES WILL ALWAYS PASS SUCH DETECTORS.” Section 31 states it “does not support using our technology to circumvent educational AI detection systems.” Refunds require “absolutely no usage…whatsoever” — zero documents processed across the account’s entire history.

A quality-tool framing is a narrower and more defensible claim than a bypass guarantee. It also sits awkwardly beside a changelog whose entries exist precisely because the output stopped passing detectors.

Has any detector company ranked it?

Read that table carefully, because it is easy to misread. The column is headed “bypassing ability on naive AI detection methods” — it rates those nine tools against weak detectors, not against GPTZero. The caption states GPTZero’s own position outright: “These methods are all ineffective against GPTZero due to the four-tiered red teaming approach.” That paper is also a GPTZero preprint, not peer-reviewed, and GPTZero has an obvious interest in the conclusion.

Phrasly is not in the appendix table of GPTZero’s February 2026 paper, which names nine bypass services individually, though GPTZero has written about it separately. It is absent from the nineteen tools in Pangram’s DAMAGE paper and from the twenty scored in Pangram’s August 2025 benchmark. We cover the Pangram side in does Phrasly bypass Pangram.

How reliable are the numbers on either side?

GPTZero’s own FAQ concedes: “Our classifier is not trained to identify AI-generated text after it has been heavily modified after generation.” Its developers page simultaneously claims a “Paraphraser Shield” means “even if AI content has been altered to look more human-like, GPTZero can detect it.” Both statements are currently published.

The two leading detector vendors publish opposite numbers on the same question. GPTZero’s February 2026 paper reports 93.5% recall on 1,000 texts run through nine bypasser services, and puts Pangram at 49.7%. Pangram’s DAMAGE paper reports GPTZero at 60.04% on humanized text, and 34.53% at GPTZero’s own default threshold. Each benchmark shows its publisher winning, and neither is peer-reviewed.

One figure to handle carefully. “GPTZero only catches 46% of humanized text” circulates constantly in humanizer marketing. It comes from Russell, Karpinska and Iyyer and it is one cell of one table: the o1-Pro humanized condition, n=30, where the “humanizer” was a research prompt the authors wrote themselves rather than any commercial tool. In the same table GPTZero scored 100% on paraphrased text and 85.3% overall at a 0.7% false positive rate. We will not quote the 46% without those qualifiers.

On false positives, GPTZero was one of seven detectors in Liang et al. (2023), which reported a 61.22% average false positive rate across those seven on 91 TOEFL essays. No per-detector figure was ever published, so that number is not GPTZero’s. GPTZero now claims it has cut its own TOEFL false positive rate to 1.1% — a figure no third party has verified.

Methodology. This page reports no first-party test of Phrasly output. On 21 September 2026 we read all fifteen entries in phrasly.ai’s public changelog, recorded each date and title, and extracted every entry mentioning GPTZero verbatim along with whether any sample size or numerical result accompanied it. We compared that against the Turnitin compatibility review of 27 August 2025 and its stated sample size. We read the current homepage FAQ and the Terms at source. We read GPTZero’s December 2024 review of Phrasly in full, recording its sample size, method, publisher and conclusion, and state its conflict of interest where it appears. Detector-side facts are quoted from GPTZero’s own FAQ, developers page and February 2026 preprint, and from Pangram’s papers and benchmark post. We have no commercial relationship with Phrasly.

What should you take from this?

Phrasly is the most transparent company in this cluster. It keeps a dated versioned changelog while its competitors publish backdated blog posts, undated guarantees and screenshots with no model version on them. That practice deserves saying out loud.

It is also the practice that produced the clearest published evidence against its own product. Three dated entries record GPTZero catching Phrasly and Phrasly patching afterwards, the detector moving first each time. There is no GPTZero pass rate anywhere, on the one detector where producing one would have been easy. And when GPTZero ran its own test, it found the output flagged as AI-paraphrased.

Transparency and a good result are not the same thing. Phrasly has the first. On the evidence it published itself, it does not reliably have the second — and you can check your own text at gptzero.me, free, no account, in about a minute.

If your real worry is being wrongly flagged on your own writing, keep your drafting history — it costs nothing and it survives a detector being wrong. It matters most for the writers detectors treat worst: ESL writers and AI detection and false positives for neurodivergent students. Check before you submit with the best pre-submission check and the Turnitin self-check, and strip the obvious tells by hand first via what to remove before using an AI humanizer.

For the product itself see our Phrasly review, and the same question against Turnitin and Pangram. Also in this cluster: StealthWriter, Rephrasy, WriteHuman.

Get the free prompt pack instead

The manual rewriting prompts that lower AI signals without a subscription — and without betting a submission on a number nobody has reproduced. Free, no card.

Grab the free prompts at detectiondrama.com →

Frequently asked questions

Does Phrasly bypass GPTZero?

Not reliably, on the vendor’s own record. Phrasly’s engineering changelog contains three dated entries in which GPTZero’s detection caught its output and Phrasly had to adapt the model in response: January 2025, June 2025 and January 2026. It publishes no GPTZero pass rate, no sample size and no screenshot. Read plainly, the changelog documents a detector repeatedly defeating a humanizer and the humanizer repeatedly patching.

What exactly do the changelog entries say?

Three entries, all repairs. On 21 January 2025: Adapted the model to GPTZero’s paraphrasing-detection update. On 17 June 2025, two detector-specific regressions were resolved, against GPTZero’s updated paraphrasing detection, which affected certain texts, and against GPTZero’s AI-editing detection. On 9 January 2026: Resolved a regression against GPTZero’s updated paraphrasing detection that affected some texts. None carries a sample size or a numerical result.

Why does the missing GPTZero review matter?

Because of which detector it is missing for. Phrasly published a dated compatibility review for Turnitin on 27 August 2025, claiming 100 documents all returning a 0% AI score. No vendor can lawfully obtain a hundred instructor-side Turnitin reports, which is the central problem with that claim. GPTZero is the opposite case: free, instant, no account, effectively unlimited. It is the one detector where a hundred-document run is genuinely achievable, and it is the one with no review at all.

Has GPTZero tested Phrasly?

Yes. On 23 December 2024 GPTZero published a hands-on review. It generated a ChatGPT essay introduction, ran it through Phrasly’s aggressive humanizer setting producing five rewrites, and scanned the results. The original scored 100% confidence as AI. After Phrasly, GPTZero flagged the output as likely AI and specifically labelled it AI paraphrased, with inconsistent results across the five drafts. Its conclusion was that both the Phrasly humanizer and Phrasly’s own detector are inconsistent and misleading.

How much weight should that carry?

Some, with caveats stated plainly. It is one source paragraph, five rewrites, published by the detector being tested, on a page arguing its own product is superior. On its own that would be weak. What gives it weight is that it points the same direction as Phrasly’s own changelog, which was written by Phrasly for engineering reasons rather than for marketing. When a vendor’s internal record and a competitor’s test agree, the finding is worth more than either alone.

Has Phrasly changed its positioning?

Noticeably. Its current homepage FAQ says the result is better writing that also scores well on platforms like Turnitin and GPTZero, not because it’s designed to trick them, but because quality writing naturally resembles human writing. It encourages users to treat Phrasly as a refinement tool, not a substitute for their own thinking, and to follow their institution’s policies. That is a quality-tool framing rather than a bypass framing, and it is a real shift.

What do its Terms say?

They disclaim the outcome in capitals: WE CANNOT GUARANTEE THAT THE TEXT GENERATED BY OUR SERVICES WILL ALWAYS PASS SUCH DETECTORS. Section 31 adds that Phrasly does not support using our technology to circumvent educational AI detection systems. Refunds require absolutely no usage whatsoever, meaning zero documents processed across the account’s entire history.

So is the changelog a point in its favour or against it?

Both, and that is the honest answer. Keeping a dated, versioned public changelog is the single most transparent practice we have found in this market, and Phrasly deserves credit for it while most competitors publish backdated blog posts and undated guarantees. The content of that changelog is also the clearest published evidence that GPTZero has caught this tool more than once. Transparency and a good result are different things, and Phrasly has the first.

About the author. Vlad Ivanov publishes Detection Drama, a site dedicated to AI-detection and text-humanization testing, with 245 published teardowns of detectors and humanizer tools including Turnitin, Pangram, GPTZero and Phrasly. Profile: LinkedIn. This page reviews published evidence and reports no first-party test; corrections with a source are welcome.