Want to bypass Turnitin in 2026? Grab the free prompt pack.
Get the exact text-humanization prompts I use to drop an AI score by hand — copy, paste, submit. Free, straight to your inbox.
Send me the free prompts →KEY TAKEAWAYS
- → less than 11 percent False positive rate the same twelve detectors show when their threshold is fitted on legacy GAN data, which is the calibration that makes them look deployable. (von Stryk and Keuper, LAION-Mobile, 2026)
- → AUC 0.624 The best score any of the twelve detectors reached on modern AI content. Five of the twelve fell below chance. (von Stryk and Keuper, LAION-Mobile, 2026)
- → 2,000 images in the WikiArt subset on which Pangram reports zero false positives. It is the benchmark the vendor uses to argue its image detector does not flag human art. (Pangram, 2026)
- → 250,000 images of historically significant artworks in the full WikiArt archive, by over 3000 artists, reaching back to the fifteenth century. A fine-art archive, not a sample of the contemporary digital illustration being flagged. (Entropy (MDPI), WikiArtVectors, 2022)
An artist runs a drawing through an AI image detector and gets a verdict back: likely AI-generated. The number looks precise. It is not a measurement of the drawing. It is the output of a model whose error rate was fixed somewhere else, on a set of pictures that probably looked nothing like the one just uploaded.
That is the finding this page is built around, and it is new. Published research covering roughly a million smartphone photographs shows that a detector’s false positive rate is not a property of the detector at all. It is a consequence of two choices made before any image is tested: where the decision threshold sits, and which pictures it was tuned against. Move either one and the same detector, on the same photographs, produces a completely different verdict rate.
Which is why the published numbers and the artist reports have never been reconcilable. The vendors are not necessarily lying. They are quoting a figure measured under conditions that no longer describe the pictures people upload. Detection Drama has tracked the same pattern on the text side, where the published false positive rates for text detectors sit far below what students and professionals report, and where detectors routinely disagree with each other on the same document.
This page collects every published false positive number for image detectors that carries a traceable source, and labels each one with the corpus it was measured on and how many images that corpus contained. The corpus labels are the part nobody else publishes, and they are where the story is.
1 How Often Do Image Detectors Flag Real Pictures as Fake?
It depends entirely on calibration. Across twelve detectors evaluated on genuine photographs, the flagging rate was less than 11 percent under the calibration those detectors are normally scored with, and 17% to 91% once the identical criterion was refit on modern content. Same detectors, same photographs, different threshold. (von Stryk and Keuper, LAION-Mobile, 2026)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Real-photo false alarm, legacy calibration | less than 11 percent | von Stryk and Keuper, LAION-Mobile |
| Real-photo false alarm, threshold refit on modern content | 17% to 91% | von Stryk and Keuper, LAION-Mobile |
| Detectors evaluated | twelve | von Stryk and Keuper, LAION-Mobile |
| Consumer-grade detectors on real photographs | 5–15% | AI Angst |
The spread is the point. A range that wide is not noise and it is not a difference between good tools and bad tools, because it is the same tools being measured twice. It is the difference between two defensible ways of deciding where likely AI begins. A vendor picking the lower figure is not fabricating anything. It is reporting the calibration that flatters its product, in a market where no convention forces it to report the other one.
For anyone on the receiving end of a flag, this collapses the distinction between a confident verdict and a hedge. A detector reporting high confidence is reporting confidence relative to a boundary somebody drew, and the paper that produced these figures found that redrawing the boundary against current imagery pushed several detectors past the point where most of the real photographs they saw came back as fake. The same structural problem shows up wherever detection is used as evidence, which is why institutional rules on using detector output as proof have been tightening rather than loosening.
It is also worth being precise about what the wide figure is and is not. It is not a claim that detectors flag most pictures most of the time in ordinary use, because most deployments run on the gentler calibration. It is a claim about how little the reassuring figure constrains real-world behaviour: a rate quoted from one calibration tells you almost nothing about performance under another, and buyers are never told which one they are getting. Treat any single published rate as a description of a test, not a promise about your file.
2 Why One Detector Reports Two Different False Positive Rates
Because the threshold is a setting, not a discovery. Detectors output a probability, and somebody chooses the cutoff at which that probability becomes an accusation. One vendor states its cutoff openly: it requires 85% before it will label an image as AI. (Sightova, 2026)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Flagging rate, threshold fitted on legacy data | less than 11 percent | von Stryk and Keuper, LAION-Mobile |
| Flagging rate, threshold refit on modern content | 17% to 91% | von Stryk and Keuper, LAION-Mobile |
| Vendor confidence required before an AI label | 85% | Sightova |
Raising the cutoff buys a lower false positive rate and pays for it in missed AI images. Lowering it does the reverse. There is no setting that gives both, and the research is explicit that none of the detectors evaluated managed to beat chance on modern AI content while keeping a defensible rate of false alarms on real ones. That is not a tuning problem waiting to be solved. It is the shape of the trade-off at the current state of the art.
The mechanism behind it matters for artists specifically. Modern phones do not simply record light; their processing pipelines fuse multiple sensor reads and suppress noise and motion blur, so the device generates the photograph as much as it captures it. Detectors read that generated quality as synthesis. Digital art tooling performs the same class of processing through brush engines, filters and upscalers, which makes the extension plausible, though it has not been measured on artwork and this page does not claim it has. Readers who want the text-side version of the same argument will find it in the work on how detection accuracy collapses outside the conditions a model was tuned for.
3 What Is Actually Inside a Vendor's Art Benchmark?
Historical paintings, mostly. The strongest vendor claim on human art reports zero false positives across a WikiArt subset of 2,000 images. WikiArt holds 250,000 works by thousands of artists, reaching back to the fifteenth century. It is a fine-art archive, not a sample of contemporary digital illustration. (Entropy (MDPI), WikiArtVectors, 2022)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Vendor art benchmark size (zero false positives reported) | 2,000 | Pangram |
| Vendor FPR, pre-2022 web images | 0.16% FPR on 10,000 images | Pangram |
| Vendor cross-detector benchmark size | 1,130 images (500 human, 630 AI) | Pangram |
| Works in the full WikiArt archive | 250,000 | Entropy (MDPI), WikiArtVectors |
A photographed oil painting and a finished digital illustration are different objects to a pixel classifier. One arrives through a camera pointed at canvas, carrying brushwork, craquelure and the texture of a physical surface. The other is born in software, built from layers and synthetic brushes, exported through a compression pipeline. A benchmark built entirely from the first tells you very little about the second, and the artists reporting problems are almost all working in the second.
The second vendor figure has the same shape. It covers pre-2022 web images, described as photographs, web graphics and computer-generated imagery, which again excludes the category under dispute. None of this makes the numbers false. It makes them answers to a question nobody was asking. The site’s own reading of that vendor’s text detector, in a review that found its near-zero false positive claim held up better than its false negatives did, reached a similar conclusion by a different route: the headline number is usually real and usually about something narrower than it sounds.
There is one more twist worth recording. The same vendor also ran an evaluation against images drawn from a subreddit where artists gather to ask whether a picture is AI, and reports classifying every human sample in it correctly. That community is simultaneously one of the places where false positive complaints are accumulating fastest. Both things can be true, and the likeliest explanation is again the sample: images that reached a clear community consensus are the easy cases, and the disputed uploads that generate the complaints are precisely the ones a consensus filter removes.
4 How Big Are the Independent Image Detector Benchmarks?
Very small. The two most widely cited independent comparisons ran on 5 real images and 5 AI-generated images and 11 images respectively. Neither separates hand-made artwork from photography, and neither publishes a false positive percentage at all. (WhichOneIsReal, 2026)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Most-cited independent comparison A | 5 real images and 5 AI-generated images | AIMultiple |
| Most-cited independent comparison B | 11 images | WhichOneIsReal |
| Largest vendor-run cross-detector benchmark | 1,130 images (500 human, 630 AI) | Pangram |
A comparison run on a handful of images cannot produce a false positive rate in any meaningful sense, and to their credit neither of these publishes one. What they publish is a ranking, and rankings from samples that small are close to arbitrary; swapping a single image can reorder the table. Those rankings are nonetheless what most roundups, buying guides and forum recommendations are ultimately built on, several citations downstream from the original.
This is the real asymmetry in the market. The vendor benchmarks are large and self-administered. The independent benchmarks are honest and tiny. The only large independent evaluation is academic, measures photographs rather than art, and reaches a conclusion none of the buying guides repeat. Anyone deciding whether to trust a verdict is choosing between those three kinds of evidence, usually without being told which one they are looking at, much as students choosing a pre-submission checker are rarely told how thin the evidence behind most detector accuracy claims actually is.
5 What a Flag Actually Tells You About Your Artwork
Less than it appears to. Human judgement is a poor benchmark here too: non-artists distinguishing human art from AI images scored 59%, with errors running in both directions at similar rates. A detector verdict is one more uncertain reading, not an adjudication. (University of Chicago, Organic or Diffused, 2024)
| METRIC | VALUE | SOURCE |
|---|---|---|
| Non-artist accuracy distinguishing human art from AI | 59% | University of Chicago, Organic or Diffused |
| Consumer-grade detector false positives on real photographs | 5–15% | AI Angst |
| Detector flagging rate once refit on modern content | 17% to 91% | von Stryk and Keuper, LAION-Mobile |
There is also a second detection channel that gets conflated with the first and works nothing like it. Platform labels such as the made-with-AI badges applied by large social networks read provenance metadata embedded in a file, while detectors classify pixels. Metadata travels with a file and can be inherited from a brush pack, a filter or an editing tool, which means a wholly handmade image can carry a signature it never earned. It can also be stripped, which is why the people generating images spend their time removing it. The two systems fail in opposite directions, and the same confusion surrounds the watermark checkers that claim to read signals that are not in the file.
Publishing platforms, galleries and marketplaces increasingly act on these verdicts, which raises the cost of the error well beyond a bruised ego. An artist who cannot clear a flag can lose a commission, a listing or an account, and the appeal process usually involves proving a negative to someone holding a number they believe is objective. The institutions furthest along this road have started backing away from that posture, which is the context behind the universities that switched their detectors off rather than keep adjudicating with them.
The practical reading is unglamorous and reasonably firm. A single verdict from a single detector is weak evidence about any individual image, whatever confidence percentage sits next to it. Agreement across several detectors is worth more, though less than it looks, because detectors share training data and tend to share mistakes. What holds up is process evidence: layered working files, version history, timestamped drafts, the messy intermediate states that a generated image does not have. That is the same advice that survives on the text side, where the first move after an accusation is to assemble the record rather than argue with the tool, and where the groups who lose most from detector error are consistently the ones whose ordinary output sits closest to whatever the model calls machine-like.

Explore every figure in this article

Methodology
Every figure on this page was taken from a named source, and each source URL was fetched and checked against the value quoted. Sources that block automated fetches, including the MDPI journal page behind the WikiArt description and the AIMultiple benchmark, were opened and read directly instead; the AIMultiple sample size and update date were both corrected from that read. Four limitations are worth stating plainly. First, no controlled study measures image detector false positive rates on contemporary hand-made digital artwork specifically, so the strongest evidence here is drawn from smartphone photography and its extension to art tooling is reasoned rather than measured. Second, artist reports about line art, finish level and brush packs are anecdotal and appear here as reported experience, never as rates. Third, vendor figures are self-administered and are labelled as such throughout. Fourth, this research ran with two stages degraded: the search-results cross-check was performed by hand rather than through the scripted route, and the backlink format filter did not run at all, so the judgement that a sourced statistics page is the right shape for this topic rests on the site’s own link history rather than on a fresh competitor analysis. The recency sweep also ran in a reduced mode, and one of its legs returned unrelated results and was discarded rather than used.
- Sources consulted: 34
- Sources cited: 8
- Data freshness: current year: 6, last year: 0, older: 2
- Data range: 2022-08-23 to 2026-09-25
- Research date: 2026-09-25
- Update schedule: Quarterly, or whenever a detector publishes a new image model
- Limitations: No controlled study measures image detector false positive rates on contemporary hand-made digital artwork specifically. The strongest evidence, LAION-Mobile, measures smartphone photographs; its mechanism is argued to extend to art tooling but that extension is reasoned, not measured, and the article must label it as such. Artist reports of line art, finish level and brush-pack effects are anecdotal and are reported as such, never as rates.
Frequently Asked Questions
How often do AI image detectors get it wrong on real images?
It depends entirely on where the threshold is set. In the LAION-Mobile evaluation of twelve detectors, the same detectors on the same photographs flagged under 11% of genuine images on their legacy calibration and 17% to 91% once the threshold was refit on modern content. (von Stryk and Keuper, LAION-Mobile, 2026)
Why do vendors report near-zero false positive rates on art?
Because of the corpus. Pangram reports zero false positives on 2,000 WikiArt images, and WikiArt is an archive of over 250,000 historically significant artworks by more than 3,000 artists, photographed. It contains almost none of the contemporary digital illustration that artists report being flagged. (Pangram and Entropy (MDPI), 2026)
How large are the independent image detector benchmarks?
Very small. The two most cited independent comparisons ran on about 10 images (AIMultiple, 14 May 2026) and 11 images (WhichOneIsReal, 3 September 2026). Neither separates hand-made artwork from photography. (AIMultiple and WhichOneIsReal, 2026)
Can a detector flag my drawing because of the software I used?
It is plausible and it is what the LAION-Mobile mechanism predicts. That paper shows smartphones with neural image processing generate rather than record photographs, which is why detectors flag them. Digital art tooling, brush engines, filters and upscalers, does the same kind of processing. Artists report exactly this, though no controlled study has measured it on art. (von Stryk and Keuper, LAION-Mobile, 2026)
Is a Made with AI label the same as a detector verdict?
No. Platform labels read provenance metadata such as C2PA and SynthID, which travels with the file, while detectors classify pixels. The two fail in opposite directions, which is why an artist can be labelled by one and cleared by the other. (Pangram, 2026)
Sources & References
- von Stryk and Keuper. “LAION-Mobile: Evaluating Deepfake Detectors On One Million Smartphone Photos.” arxiv.org/abs/2609.11134. Accessed 2026-09-25.
- Pangram. “Introducing Pangram Image Detection (Research Preview).” pangram.com/blog/introducing-pangram-image-detection. Accessed 2026-09-25.
- Entropy. “WikiArtVectors: Style and Color Representations of Artworks.” mdpi.com/1099-4300/24/9/1175. Accessed 2026-09-25.
- AI Angst. “AI Image Detectors: What Actually Works in 2026, and What Doesn't.” aiangst.com/review/ai-image-detectors. Accessed 2026-09-25.
- WhichOneIsReal. “Best AI Image Detector Tools in 2026 (Free & Paid, Compared).” whichoneisreal.com/compare/best-ai-image-detector/. Accessed 2026-09-25.
- AIMultiple. “AI Image Detector Benchmark.” aimultiple.com/ai-image-detector. Accessed 2026-09-25.
- Sightova. “Testing Methodology.” sightova.com/testing-methodology. Accessed 2026-09-25.
- University of Chicago. “Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?.” arxiv.org/html/2402.03214v1. Accessed 2026-09-25.
Last updated:
