TopAIHumanizer › Guide
Do AI detectors actually work?
Not reliably — and the strongest evidence for that comes from the detector vendors themselves. OpenAI withdrew its own classifier for poor accuracy. Turnitin suppresses its scores below 20% and tells instructors not to act on them alone. Independent research found error rates that would be disqualifying in any other assessment tool.

The published evidence
Stanford: 61% false positive rate on non-native writers
Researchers tested seven AI detectors against writing by non-native English speakers. The detectors misclassified genuine human writing as AI-generated 61% of the time. On roughly 20% of papers, all seven detectors were wrong simultaneously. The same tools almost never made this error on native speakers' work.
The cause is structural, not a bug that will be patched. Detectors flag low lexical diversity and predictable sentence construction. Competent non-native writing frequently exhibits both. The bias is baked into what these models measure.
OpenAI withdrew its own detector
In July 2023 OpenAI discontinued its AI Text Classifier, citing low accuracy. The organisation with the most direct knowledge of how its own models generate text concluded it could not reliably identify that text after the fact.
Turnitin's own caveats
Turnitin claims a false positive rate under 1%, but only for documents already scoring above 20% AI writing — a conditional that is easy to miss. More telling: it suppresses all scores between 1% and 19% entirely, displaying an asterisk. It requires 300+ words before producing any score. And its documentation states the indicator "should not be used as the sole basis for action."
Quill.org and CommonLit discontinued their AI Writing Check tool. A number of universities have disabled Turnitin's AI indicator institution-wide.
What this adds up to: these tools produce a number that feels precise and is not. The precision is the dangerous part — a "94% AI" readout invites confidence that the underlying method does not support.
Why detection is hard in principle
Detectors measure statistical properties of text — how predictable each word is given the preceding ones (perplexity), and how much that predictability varies (burstiness). Machine-generated text tends to sit in a narrower band on both.
The problem is that plenty of human writing does too. Technical documentation. Legal drafting. Writing by anyone taught to write plainly and consistently. Non-native writing. Heavily edited prose. All of it looks statistically similar to model output, because clarity and predictability are correlated.
Conversely, asking a model to vary its sentence length moves it out of the flagged band immediately. The signal is easy to evade deliberately and easy to trip accidentally — the worst combination for an assessment tool.
What a score actually tells you
- A high score is weak evidence, not proof. It is consistent with AI use. It is also consistent with being a careful writer, a non-native speaker, or someone writing in a formulaic genre.
- A low score proves nothing at all. Trivial edits move scores substantially. A low score cannot establish human authorship.
- Scores are not comparable across tools. GPTZero, Originality.ai, Copyleaks and Turnitin routinely disagree on the same document, often dramatically.
- Scores are not stable over time. Models update. The same document can score differently in March and September.
What is replacing detection
Institutions that have thought carefully about this are moving to process evidence rather than artifact analysis: version history, draft submissions at intervals, in-person defence of submitted work, and disclosure requirements. These are harder to fake and do not carry a demographic bias.
For writers, the practical implication is simple and cheap: keep your version history. Google Docs and Word retain it automatically. It is far stronger protection against a false accusation than any argument about a percentage.
Our position
We compare AI humanizers, so it would be commercially convenient for us to tell you detectors are all-powerful and you need a tool to beat them. The evidence does not support that, and we are not going to pretend otherwise.
Detectors are unreliable enough that optimising against them is a poor use of effort. Write clearly, keep your drafts, disclose where asked. If you use a humanizer, use it because the writing is genuinely clunky — not because a percentage frightened you.
Sources: The Markup — AI detection tools falsely accuse international students · Turnitin AI writing detection FAQs
FAQ
Common questions.
How accurate are AI detectors?
Published research found seven detectors misclassified non-native English speakers' genuine writing as AI-generated 61% of the time. Vendor-claimed accuracy figures are typically conditional and self-measured.
Why do AI detectors flag human writing?
They measure text predictability and variation. Clear, consistent human writing — including most non-native writing and technical prose — shares those statistical properties with model output.
Did OpenAI shut down its AI detector?
Yes. OpenAI discontinued its AI Text Classifier in July 2023, citing low accuracy.
Do different AI detectors agree with each other?
Frequently not. GPTZero, Originality.ai, Copyleaks and Turnitin often return substantially different scores for the same document.
What should I do if I'm falsely flagged?
Produce process evidence — version history, drafts, research notes. Cite the detector vendor's own guidance on score limitations and the published false-positive research. Version history is the strongest single protection.