What Actually Survives Humanization? Every Published Number on Which AI Writing Markers Editing Removes (2026)
Strip out every surface-level AI artifact a humanizer targets — the clichés, the purple prose, the redundant exposition — and a structure-based detector’s accuracy falls from 95.5 to 93.9 macro-F1. A 1.6-point drop. The signals humanizers rewrite are not the signals that identify the text.
Source: StoryScope, COLM 2026 (arXiv:2604.03136v6)
Want to bypass Turnitin in 2026? Grab the free prompt pack.
Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.
Send me the free prompts →- 1.6 points — the entire accuracy cost of removing AI writing artifacts span by span, measured on 278 rewritten stories (StoryScope, 2026).
- 97% → 3% — by contrast, what happens to style-based detection when a model is fine-tuned to imitate human prose (Chakrabarty et al. 2026, cited in StoryScope).
- 93.2 macro-F1 — human-versus-AI detection using narrative structure alone, with every stylistic signal stripped out. Only 2.8 points below the full model (StoryScope, 2026).
- 55.9 macro-F1 — zero-shot perplexity detection on the same corpus, barely above a coin flip. This is the signal class most consumer detectors run on (StoryScope, 2026).
- 81% vs 38% — how often AI renders emotion as physical sensation rather than naming it, against human authors. The largest single marker gap measured (StoryScope, 2026).
- Severe degradation — what happens to native-language traces once fluency editing and paraphrasing are applied to 450 essays. Not every marker is durable (arXiv:2605.10216, 2026).
- 61,608 stories — corpus size: 10,272 prompts, each written by a human author and five separate LLMs, 304 features extracted per story (StoryScope, 2026).
1What editing actually removes — and what it leaves
Editing removes the surface. It does not remove the structure. In the largest controlled test published to date, researchers took AI-generated stories, rewrote them span by span to strip out seven categories of recognisable AI writing artifacts, and then re-ran a detector built purely on narrative structure. Accuracy fell 1.6 points.
That number is the whole argument. Humanizers operate on word choice, sentence rhythm, and phrasing. Those are real signals, and they are genuinely fragile — which is exactly why humanized text posts such different scores across detectors. But a second class of signal sits underneath the prose and does not move when the prose does.
| Metric | Value | Source |
|---|---|---|
| Corpus size | 61,608 stories | StoryScope, COLM 2026 |
| Prompts, each written 6 ways | 10,272 | StoryScope, COLM 2026 |
| Features extracted per story | 304 | StoryScope, COLM 2026 |
| Mean story length | 4,753 words | StoryScope, COLM 2026 |
| Stories in the editing test | 278 | StoryScope §4.2 |
| Accuracy before editing | 95.5 macro-F1 | StoryScope §4.2 |
| Accuracy after editing | 93.9 macro-F1 | StoryScope §4.2 |
Compare that to what happens on the other side of the line. The same paper cites work showing that fine-tuning a model to mimic human style drops detection on creative writing from 97% to 3%. Surface style can be defeated almost completely. Narrative structure lost 1.6 points to a dedicated rewriting pass. Those two numbers describe two entirely different kinds of evidence, which is a large part of why detectors disagree so violently with each other on the same document.
2Detection accuracy by signal class
Not all detection methods are looking at the same thing. Ranked by how they performed on this corpus, the spread is wide — and the method most consumer tools rely on finished last.
| Method | Macro-F1 | Source |
|---|---|---|
| Narrative + style (304 features) | 96.0 | StoryScope Table 2 |
| Edited text, artifacts removed | 93.9 | StoryScope Table 2 |
| Narrative structure only (257 features) | 93.2 | StoryScope Table 2 |
| Style signals only (39 features) | 85.8 | StoryScope Table 2 |
| 30 core narrative features | 84.8 | StoryScope Table 2 |
| Binoculars (zero-shot perplexity) | 55.9 | StoryScope Table 2 |

One caveat the researchers are candid about: raw text-based baselines beat all of these on unedited text, with stylometric and TF-IDF classifiers scoring above 99.5. The narrative model’s claim is not that it is the most accurate detector available. It is that its accuracy is the kind that survives someone trying to remove it.
3The specific markers that give AI away
The structural signal is not abstract. It decomposes into measurable narrative habits, each present at a different rate in AI and human writing. The largest gap is emotional: AI writes the body, humans name the feeling.
| Marker | AI vs Human | Source |
|---|---|---|
| Emotion shown as physical sensation | 81% / 38% | StoryScope Table 16 |
| Olfactory sensory detail | 82% / 57% | StoryScope Table 16 |
| No subplots at all | 79% / 57% | StoryScope Table 16 |
| Narrator comments on the theme | 77% / 52% | StoryScope Table 16 |
| Resolution driven by protagonist choice | 69% / 46% | StoryScope Table 16 |
| Dialogue as philosophical debate | 59% / 34% | StoryScope Table 16 |
| Fourth-wall permeability | 39% / 67% | StoryScope Table 16 |
| Direct reader address | 7% / 28% | StoryScope Table 16 |
| Explicit named references to other works | 24% / 47% | StoryScope Table 16 |
| Ambivalent or mixed moral position | 38% / 59% | StoryScope Table 16 |
Read as a group, these describe a writer that over-explains. AI states the theme, resolves the moral question, gives the protagonist agency over the ending, and cuts the subplots. Humans leave things ambivalent, break the fourth wall, and reference the world outside the text. It is a structural version of the same instinct that makes writing feel machine-made to an ordinary reader long before any detector is involved — and it is not something a pre-rewrite cleanup pass can address, because none of it lives at the sentence level.

4The markers that rewriting does erase
Durability is not universal, and the counter-example is well documented. A separate 2026 study ran 450 essays from the Write & Improve corpus through escalating levels of intervention to see how long native-language traces survive.
| Intervention level | Effect on L1 traces | Source |
|---|---|---|
| Minimal grammar correction | Traces preserved | arXiv:2605.10216 |
| Fluency editing | Severe degradation | arXiv:2605.10216 |
| Full paraphrase | Severe degradation | arXiv:2605.10216 |
The mechanism is instructive. Native-language profiling does not run on spelling errors; it runs on unidiomatic lexico-semantic choices and pragmatic transfer. Those are deeper than grammar and shallower than structure — and paraphrasing normalises them away. This is the same underlying property that drives the documented bias of AI detectors against non-native English writers: the features that mark a second-language writer sit in precisely the layer that rewriting flattens.
So the honest summary is a split, not a verdict. Lexical and stylistic fingerprints — vocabulary, idiom, rhythm, punctuation habits — are the fragile class. Discourse-level construction is the durable class. Any claim that humanizing “works” or “doesn’t work” is answering the wrong question until it says which class it means, a distinction largely missing from the standard explanation of how humanizers operate.
5Each model leaves a different fingerprint
Beyond the human-versus-AI line, the five models tested were separable from each other on narrative features alone — 68.4% accuracy across six classes against a 16.7% chance baseline.
| Model | Distinguishing habit | Source |
|---|---|---|
| Gemini 3 Flash | 88% bleak settings | StoryScope fingerprints |
| GPT-5.4 | 64% gossip-driven plots | StoryScope fingerprints |
| Claude Sonnet 4.6 | 62% reverent to tradition | StoryScope fingerprints |
| Kimi K2.5 | 3 fingerprint features | StoryScope fingerprints |
| Human authors | 32 fingerprint features | StoryScope fingerprints |
Human stories were also measurably stranger: 24.7% landed in the rarest tenth of the corpus against 7.1% of AI stories. The practical read is that models have converged on a shared way of telling stories, and the distance between any two of them is smaller than the distance between all of them and people. That convergence is what makes structural detection viable at all, and it is the same trend behind detectors learning to recognise humanizer output as its own category.
Which markers is your edit actually touching?
Select the deepest level of change made to the text. The result reflects the evidence classes above — it is not a detector score or a prediction about any specific tool.
6What this does not prove
The headline finding rests on a narrower base than the number suggests, and the limits belong in the open.
The editing test used 278 stories, one rewriting framework, and one model rewriting its own output. There is no published test of a different model paraphrasing, no human-editor rewrite, and no adversarial attempt to change narrative structure directly.
The corpus is roughly 5,000-word fiction. Features like subplot integration and chronological discontinuity do not exist in a 600-word discussion post. Applying these numbers to coursework, applications, or marketing copy is extrapolation, not evidence — a caution worth holding alongside what commercial detectors are actually measuring.
None of this is a commercial detector. The models here are research classifiers trained on extracted narrative features. No mainstream tool ships this method today, so a strong structural signal in a paper does not mean the detector reading your document can see it.
What the evidence does support is narrower and still useful: the durability of an authorship signal depends on which layer it lives in, and the layer humanizers operate on is the shallow one. That gap between how widely these tools are now used and what they structurally can and cannot change is the part of the picture that gets least attention.
Methodology
Every figure on this page is drawn from two peer-reviewed or preprint sources published in 2026, listed in full below. Data was compiled on 14 August 2026 from the primary papers rather than secondary coverage; tables 2, 3 and 16 and Section 4.2 of the StoryScope paper supplied the detection and marker figures, and arXiv:2605.10216 supplied the editorial-intervention findings.
Freshness: both sources are from 2026. No figure on this page predates 2026, and the 97%-to-3% style-detection figure is reported as prior work cited within StoryScope rather than as an original StoryScope measurement.
Limitations: the StoryScope paper contains no formal limitations section; the constraints described in section 6 above are drawn from statements the authors make in the body of the paper. Split counts reported in the paper are internally inconsistent between sections 3 and D, so no split figure is quoted here. This page will be updated when replication on non-fiction text is published.
Frequently asked questions
Does humanizing AI text actually remove the AI markers?
It removes one class and leaves another. Span-level removal of AI writing artifacts dropped structural detection accuracy by only 1.6 macro-F1 points (95.5 to 93.9) in the 2026 StoryScope study, while separate research found style-mimicking fine-tuning can collapse style-based detection from 97% to 3%. Surface style is defeatable; the underlying structural choices largely are not.
Which AI writing markers survive editing?
Discourse-level choices: how causally linear the narrative is, how explicitly the theme is stated, whether emotion is shown as physical sensation, and whether subplots exist at all. StoryScope reports narrative-only detection at 93.2 macro-F1 with zero stylistic signal. Changing these requires structural rewriting rather than word substitution.
Which markers do get erased by rewriting?
Surface and lexical ones. A 2026 study of 450 Write & Improve essays found native-language traces survive minimal grammar correction but suffer severe degradation once fluency edits and paraphrasing normalise lexico-semantic choices. Vocabulary-level fingerprints are the fragile class.
Do these findings apply to student essays?
Not directly, and the limit matters. The StoryScope corpus is roughly 5,000-word fiction, and the editing-robustness result rests on 278 stories rewritten by a single model. No published test extends it to essays, reports, or marketing copy, so applying the numbers to coursework is extrapolation rather than evidence.
Is counting em-dashes a real AI tell?
It is the fragile kind. StoryScope notes that GPT-5.4 significantly reduced em-dash usage relative to earlier models, which is exactly how surface signatures behave: they shift with each model release. Structural markers moved far less, which is why the researchers argue they offer a more durable basis for authorship analysis.
Can detectors tell which AI model wrote something?
Partially. Using narrative features alone, six-way attribution across human plus five models reached 68.4% accuracy against a 16.7% chance baseline. The confusions are almost entirely AI-versus-AI, meaning the boundary between human and machine is far sharper than the boundaries between individual models.
Why do commercial AI detectors miss text these methods catch?
Different signal class. Binoculars, a zero-shot perplexity method of the kind most consumer detectors rely on, scored 55.9 macro-F1 on this corpus, barely above chance. Perplexity is precisely what humanizers are built to manipulate, and it is a large part of why detection results vary so much between tools.
How different is AI writing from human writing overall?
The five models cluster together and away from people. Mean human-to-AI distance in feature space was 1.6 times the mean AI-to-AI distance (6.6 versus 4.3), and the closest human-AI pair still sat farther apart than the most distant AI-AI pair. Human stories were also rarer: 24.7% fell in the corpus’s rarest tenth against 7.1% of AI stories.
Sources
- Russell, J., Rajendhran, R., Pham, C. M., Wieting, J., & Iyyer, M. “StoryScope: Investigating idiosyncrasies in AI fiction.” COLM 2026, arXiv:2604.03136v6. arxiv.org. Accessed 14 August 2026.
- StoryScope full text — Tables 2, 3, 16 and Section 4.2. arxiv.org (PDF). Accessed 14 August 2026.
- StoryScope code and released corpus (51,336 AI-generated narratives across 10,272 prompts). github.com. Accessed 14 August 2026.
- “The Impact of Editorial Intervention on Detecting Native Language Traces.” arXiv:2605.10216. arxiv.org. Accessed 14 August 2026.
Last updated: August 14, 2026. Compiled by the Detection Drama Research Team.
