# Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?

Source: https://forum.global100.org/q/why-do-gptzero-zerogpt-quillbot-and-grammarly-give-different-ai-scores/
Site: Global 100 Forum, category Detector tools compared
Published: 2026-09-17. Updated: 2026-09-23. Replies: 1.

## Question by Maya Lindqvist (staff), 2026-09-17

They disagree because they are answering different questions with different models, different training data and different cut-offs. A "70%" from one tool is a probability, from another a share of sentences, and from a third a share of text. The caveat: none of them measure the same thing, so averaging or comparing the numbers is meaningless.

## Two families of method

The first family measures how predictable the text is to a language model. Perplexity is the model's surprise at each word; burstiness is how much that surprise varies across sentences. GPTZero's original 2023 model worked this way, and its [own explainer](https://gptzero.me/news/perplexity-and-burstiness-what-is-it/) said a perplexity above about 85 was more likely human. The same page now carries an update: as of autumn 2023 GPTZero no longer uses perplexity and burstiness, having moved to a deep learning classifier. A more refined statistical approach, [DetectGPT by Mitchell and colleagues (ICML 2023)](https://arxiv.org/abs/2301.11305), looks at the curvature of the model's probability function rather than raw perplexity, and needs no training data at all.

The second family trains a classifier on large sets of labelled human and AI text and lets it learn whatever patterns separate them. Turnitin's [FAQ](https://guides.turnitin.com/hc/en-us/articles/28477544839821-Turnitin-s-AI-writing-detection-capabilities-FAQs) states plainly that its model is not programmed to evaluate burstiness or perplexity and that individual predictions may not be explainable feature by feature. GPTZero, [Grammarly](https://www.grammarly.com/ai-detector) and [QuillBot](https://quillbot.com/ai-content-detector) all describe their current detectors as trained models of this kind, and QuillBot says its model still weighs predictability, sentence variation and repetitiveness among its signals.

## Training data sets the baseline

A classifier only knows what it was shown. Grammarly's page says its model was trained on tens of thousands of texts. GPTZero says millions of documents across genres. [ZeroGPT](https://www.zerogpt.com/) describes web, educational and proprietary synthetic datasets. If a tool saw little formal academic prose, or few non-native writers, it will score those inputs differently from a tool that did. The [bias thread](/q/are-ai-detectors-biased-against-non-native-english-writers/) shows how far that can go.

## Thresholds decide what a number becomes

Every tool ends with a probability and then has to turn it into a label. Where it draws the line is a business decision, not a fact about the text. QuillBot says that when a result is unclear its model leans towards "human" to reduce false positives. GPTZero says its output mapping is deliberately biased towards false negatives over false positives and that its "high confidence" band is tuned to under 1% error. Two tools with identical raw scores can therefore show different verdicts.

**What each tool's score actually represents**

| Tool | Method (as described by the vendor) | What the score means | What a "70%" means here |
| --- | --- | --- | --- |
| GPTZero | Sentence-by-sentence deep learning classifier; earlier versions used perplexity and burstiness | Document class (human, mixed, AI) with class probabilities and a confidence band | 70% probability of the AI class, read alongside the confidence label, not a share of the text |
| ZeroGPT | Multi-stage deep learning model the vendor calls DeepAnalyse; claims 98.4% accuracy and under 1% false positives | Percentage of the text flagged as AI, with highlighted sentences | About 70% of the text was flagged, judged sentence by sentence |
| QuillBot | Trained language-model classifier weighing predictability, sentence variation and repetition; 80-word minimum | 0 to 100% likelihood the text was AI-generated, plus sentence highlights and AI-refined labels | 70% likelihood, with uncertain cases pushed towards human |
| Grammarly | Segment-by-segment trained classifier | Percentage of the text that "appears to be AI-generated" | About 70% of segments matched AI patterns |

## What that means in practice

- **Compare labels, not numbers.** A 70% from GPTZero and a 70% from ZeroGPT are different units. Only the tool's own documentation tells you which.
- **Short text inflates disagreement.** QuillBot requires 80 words and Turnitin 300; below that, a single sentence can swing the score.
- **Benchmarks exist but vendors pick them.** The [RAID benchmark (Dugan and colleagues, ACL 2024)](https://arxiv.org/abs/2405.07940) found detectors were easily fooled by sampling changes and adversarial edits, which is why a [watermark](/q/does-watermarking-ai-generated-text-actually-work/) or a writing record is stronger evidence than any score.

## Reply 1 by Maya Lindqvist (staff), 2026-09-23

If you want to see the threshold effect for yourself, run one long human-written paragraph through two tools, then delete the most formulaic sentence and run it again. The percentage tools (ZeroGPT, Grammarly) tend to move in small steps because one sentence is a small share of the text. The probability tools (GPTZero, QuillBot) can jump from one label to another when a borderline document crosses the cut-off. Neither behaviour is wrong; it just shows that the number is downstream of a decision the vendor made, not a property of your writing.

---
Cite as: Global 100 Forum, "Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?", https://forum.global100.org/q/why-do-gptzero-zerogpt-quillbot-and-grammarly-give-different-ai-scores/, accessed 2026-10-11.
