How accurate are AI text detectors, really?
Less accurate than their marketing suggests, and far less reliable on edited or paraphrased text. Independent tests find that detectors catch unedited chatbot output reasonably often, miss much of it once a person or a tool rewrites it, and occasionally flag writing that a human produced. A score is a signal to look closer, not a finding.
Cite
Global 100 Forum, "How accurate are AI text detectors, really?", https://forum.global100.org/q/how-accurate-are-ai-text-detectors-really/, accessed 2026-10-11.- Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·
The most useful independent study is still Weber-Wulff and colleagues' 2023 test of 14 detection tools, published in the International Journal for Educational Integrity. The team wrote human texts, generated texts with ChatGPT, and then created harder cases: machine-translated text and text that had been manually edited or paraphrased. Their conclusion was blunt. None of the tools was accurate or reliable enough to be used as evidence, and all of them struggled once the AI text had been obfuscated.
The strongest signal of all came from a detector's own maker. OpenAI launched an AI text classifier in January 2023 and withdrew it that July because of its low rate of accuracy. In OpenAI's own evaluation it labelled only 26% of AI-written text as "likely AI-written" and wrongly flagged 9% of human text.
Key figures from the sources above
Source What was tested Result OpenAI classifier, 2023 OpenAI's own evaluation of its detector 26% of AI text caught, 9% of human text wrongly flagged; withdrawn July 2023 Weber-Wulff et al., 2023 (see also the humanizers thread) 14 detection tools on human, AI, translated and paraphrased text No tool judged accurate or reliable enough for evidence; all degraded on obfuscated text Turnitin, 2023 (as cited by Vanderbilt) Vendor's stated false positive rate 1% claimed, which Vanderbilt estimated at about 750 wrongly flagged papers a year at its volume Two related threads cover the parts of this question that come up most: whether detectors are biased against non-native English writers and what a student should do when a detector flags their work.
How to read a score
- Treat the number as a probability estimate about the text, not a statement about the writer. Two documents with the same score can have very different histories.
- Short texts are unreliable. Most tools say so in their documentation. A paragraph or a short answer does not give a model enough to go on.
- Editing moves scores a lot. Light human editing, grammar tools and translation can all push a score in either direction.
- Look for corroborating evidence. Drafts, version history, notes and a conversation with the writer tell you far more than a percentage.
- Prefer provenance where it exists. Signals planted by the producer, such as text watermarks, are stronger evidence than a style guess, though only when the provider added them.
1 more reply
Most helpful first- Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·
A follow-up, because this comes up in almost every thread about scores: the tools have changed since 2023, so it is fair to ask whether the older studies still apply.
Some things have improved. Several vendors now report sentence-level highlighting and publish false positive rates for long documents. What has not changed is the underlying problem. Detectors infer authorship from statistical patterns in the text, and those patterns overlap between fluent human writing and model output. Any threshold that catches more AI text also flags more human text. That trade-off is why we keep saying the same thing: use a detector to decide where to look, never to decide what happened.
Write something first.
Give people something to work with: at least 30 words on what happened and what you tried.
That is too long. Keep it under 6,000 characters.
Write the question as the title, 15 to 140 characters, no links.
Pick a category.
Add a name (2 to 40 characters, no links).
That email address does not look right.
That was quick. Read the thread, then try again.
The form expired. Reload the page and post again.
Something went wrong with the form. Reload and try again.
Please complete the check and post again.
Limit reached for now. Try again later.
This thread is closed to new replies.
This thread no longer accepts replies.
Something went wrong with the form. Reload and try again.
Post a reply or question first (name and email), then this browser can vote, edit and accept answers.
You cannot vote on your own post.
Only the author (within 30 days) or the forum team can do that.
That email belongs to a forum team account. Use your sign-in link instead.
New accounts are paused for the moment. Try again later.
That was already posted.