# How accurate are AI text detectors, really?

Source: https://forum.global100.org/q/how-accurate-are-ai-text-detectors-really/
Site: Global 100 Forum, category AI text detection
Published: 2026-09-08. Updated: 2026-09-12. Replies: 1.

## Question by Maya Lindqvist (staff), 2026-09-08

Less accurate than their marketing suggests, and far less reliable on edited or paraphrased text. Independent tests find that detectors catch unedited chatbot output reasonably often, miss much of it once a person or a tool rewrites it, and occasionally flag writing that a human produced. A score is a signal to look closer, not a finding.

The most useful independent study is still [Weber-Wulff and colleagues' 2023 test of 14 detection tools](https://link.springer.com/article/10.1007/s40979-023-00146-z), published in the International Journal for Educational Integrity. The team wrote human texts, generated texts with ChatGPT, and then created harder cases: machine-translated text and text that had been manually edited or paraphrased. Their conclusion was blunt. None of the tools was accurate or reliable enough to be used as evidence, and all of them struggled once the AI text had been obfuscated.

The strongest signal of all came from a detector's own maker. OpenAI launched an AI text classifier in January 2023 and [withdrew it that July because of its low rate of accuracy](https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/). In OpenAI's own evaluation it labelled only 26% of AI-written text as "likely AI-written" and wrongly flagged 9% of human text.

**Key figures from the sources above**

| Source | What was tested | Result |
| --- | --- | --- |
| OpenAI classifier, 2023 | OpenAI's own evaluation of its detector | 26% of AI text caught, 9% of human text wrongly flagged; withdrawn July 2023 |
| Weber-Wulff et al., 2023 (see also the [humanizers thread](/q/do-ai-humanizers-and-paraphrasers-actually-beat-ai-detectors/)) | 14 detection tools on human, AI, translated and paraphrased text | No tool judged accurate or reliable enough for evidence; all degraded on obfuscated text |
| Turnitin, 2023 (as cited by Vanderbilt) | Vendor's stated false positive rate | 1% claimed, which Vanderbilt estimated at about 750 wrongly flagged papers a year at its volume |

Two related threads cover the parts of this question that come up most: [whether detectors are biased against non-native English writers](/q/are-ai-detectors-biased-against-non-native-english-writers/) and [what a student should do when a detector flags their work](/q/what-should-a-student-do-if-an-ai-detector-flags-their-work/).

## How to read a score

- **Treat the number as a probability estimate about the text, not a statement about the writer.** Two documents with the same score can have very different histories.
- **Short texts are unreliable.** Most tools say so in their documentation. A paragraph or a short answer does not give a model enough to go on.
- **Editing moves scores a lot.** Light human editing, grammar tools and translation can all push a score in either direction.
- **Look for corroborating evidence.** Drafts, version history, notes and a conversation with the writer tell you far more than a percentage.
- **Prefer provenance where it exists.** Signals planted by the producer, such as [text watermarks](/q/does-watermarking-ai-generated-text-actually-work/), are stronger evidence than a style guess, though only when the provider added them.

## Reply 1 by Maya Lindqvist (staff), 2026-09-12

A follow-up, because this comes up in almost every thread about scores: the tools have changed since 2023, so it is fair to ask whether the older studies still apply.

Some things have improved. Several vendors now report sentence-level highlighting and publish false positive rates for long documents. What has not changed is the underlying problem. Detectors infer authorship from statistical patterns in the text, and those patterns overlap between fluent human writing and model output. Any threshold that catches more AI text also flags more human text. That trade-off is why we keep saying the same thing: use a detector to decide where to look, never to decide what happened.

---
Cite as: Global 100 Forum, "How accurate are AI text detectors, really?", https://forum.global100.org/q/how-accurate-are-ai-text-detectors-really/, accessed 2026-10-11.
