How Accurate Are AI Detectors? What Scores Really Mean
How accurate are AI detectors? On long, untouched machine-written text, the better ones get it right the large majority of the time. On short, edited, translated or deliberately reworded text, they get noticeably worse — and they rarely tell you which situation you're in.
That gap is the whole story. A detector can be highly accurate in aggregate and still be wrong about the specific document in front of you.
So the useful question isn't "which tool is right" but "what is this score actually measuring, and what would change my mind about it?"
What a detection score actually measures
An AI detector doesn't recognise ChatGPT's handwriting. It measures how closely the statistical shape of a piece of text resembles patterns it has learned to associate with model-generated writing: word choices that are unusually predictable, sentence lengths that vary less than a person's, phrasing that sits close to the average of everything the model was trained on.
Older approaches leaned on perplexity and burstiness — roughly, how surprising each next word is and how much that surprise fluctuates. Current tools mostly use trained classifiers fed millions of human and AI samples, which is why they perform better but also why they inherit the quirks of whatever data they were trained on.
Two things follow from this. A score is a similarity estimate, not a record of what happened at the keyboard. And "68% AI" does not mean 68% of the document was generated — it usually means the classifier's confidence, however that vendor chose to express it.
How accurate are AI detectors in independent testing?
Vendors publish strong numbers. Turnitin's own documentation states its AI writing detection has a false positive rate under 1% for documents it flags as 20% or more AI-written, while noting a higher rate at the individual sentence level, which is why it advises against acting on sentence highlights alone (Turnitin AI writing detection FAQ). GPTZero, Copyleaks, Winston AI, Pangram and the Grammarly AI detector all report accuracy figures in the high nineties.
Read those numbers as what they are: performance on a test set the vendor assembled, usually with clean, full-length samples and a known set of generating models.
Independent evaluations tend to be less flattering and much more varied. Rankings shift whenever a new model ships, and the spread between tools on the same corpus can be enormous. A Stanford-led study published in Patterns found that several detectors misclassified writing by non-native English speakers as AI-generated at strikingly high rates, and that simple prompting tricks could move text past them (Liang et al., 2023).
There's also a base-rate problem that headline accuracy hides. A 1% false positive rate sounds excellent until you apply it to 400 submitted essays — you should expect roughly four human-written papers to be flagged, every single time.
When scores stop being trustworthy
Accuracy isn't a fixed property of a tool. It's a property of the tool plus the text you gave it. These are the conditions that degrade it most:
- Short samples. Under roughly 150–300 words there simply isn't enough signal. Most vendors say so in their own documentation.
- Formulaic genres. Lab reports, legal summaries, product descriptions and technical instructions are repetitive and low-variance by design — which is exactly what detectors read as machine-like.
- Non-native English. Simpler vocabulary and more predictable syntax push scores up for reasons that have nothing to do with AI use.
- Heavy editing. A person rewriting AI output, or an AI polishing human prose, produces mixed text that both categories fit badly.
- Paraphrase and humanizer tools. Anything marketed to humanize AI text works by disrupting the statistical fingerprint. It often succeeds against one detector and fails against another.
- New models. A classifier trained before a model existed has never seen its output distribution.
None of this makes detection worthless. It means a score needs context before it means anything, which is the argument we make in more depth in our guide to reading an artificial intelligence detector's evidence.
Read the passages, not the percentage
The single number is the least informative part of any result. What's worth your attention is where the tool found its signal.
A document where three consecutive paragraphs light up while the rest reads clean tells you something specific and checkable. A document with a scattering of mild highlights across twenty pages tells you almost nothing.
That's why every result on this site shows passage-level evidence rather than a pass-or-fail verdict. If you want to see the difference on your own writing, run a document through the text detector and look at which sentences it marks, then ask yourself whether you can explain each one. For coursework specifically, the essay detector is set up for longer submissions, and the ChatGPT detector is worth a second pass when you suspect a particular model.
Images follow the same logic
An AI image detector is probabilistic too, and for similar reasons: it reads compression artefacts, texture statistics and generator fingerprints. Screenshots, re-saves, filters and heavy cropping all strip the signal it depends on.
What to do when a score comes back high
Treat a high score as a prompt to look for evidence a classifier can't produce.
- Check length and genre first. If the passage is 80 words of methods section, stop — the result isn't meaningful.
- Run it again with a second tool. Agreement between independent detectors is more informative than one confident number. Disagreement is itself a finding.
- Look at process artefacts. Google Docs version history, editor revision trails, drafts, notes and timestamps are far harder to fake than prose style.
- Compare against known work. A sudden jump in register from the same author is a stronger signal than any percentage.
- Ask about the content. A five-minute conversation about the argument and its sources settles most cases faster than any checker.
What you should never do is confront someone with a score as if it were proof. Detectors estimate resemblance; they do not establish authorship, and treating them otherwise is how institutions end up apologising.
Frequently asked questions
Do AI detectors work at all, or are they just guessing?
They work in the sense that they perform far better than chance on long, unmodified AI text — that's a measurable, repeatable result. They're not guessing, but they are estimating, and the estimate gets shakier as text gets shorter, more edited or more formulaic. Useful signal, not a verdict.
What are the best AI detectors right now?
No single tool wins across every condition, and any ranking you read will be out of date within a few model releases. Turnitin has the most institutional data behind it, GPTZero and Copyleaks are widely used in education, and Pangram and Winston AI perform well in some independent benchmarks. Running two and comparing the highlighted passages beats picking a favourite.
Can an AI detector be wrong about my own writing?
Yes, and it happens most often to careful, plain, well-structured writers — plus anyone writing in English as a second language. If your work is flagged, your best defence is process evidence: version history, drafts, research notes and a willingness to talk through how the piece came together.
Do humanizer tools beat AI detectors?
Sometimes, against some detectors, temporarily. Humanizers add stylistic noise that can lower a score, but they also tend to introduce odd word choices that a human reader notices, and detector vendors retrain against these tools continuously. It's an arms race with no stable winner.
One thing to do differently
Next time a score comes back high, write down what else you'd need to see before you'd act on it — a draft, a version history, an answer to a question about paragraph three. If you can't name that second piece of evidence, you don't have a case yet; you have a probability. If you've hit a result you can't interpret, send it to us and we'll tell you what we'd look at.