Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?
They disagree because they are answering different questions with different models, different training data and different cut-offs. A "70%" from one tool is a probability, from another a share of sentences, and from a third a share of text. The caveat: none of them measure the same thing, so averaging or comparing the numbers is meaningless.
Cite
Global 100 Forum, "Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?", https://forum.global100.org/q/why-do-gptzero-zerogpt-quillbot-and-grammarly-give-different-ai-scores/, accessed 2026-10-11.- Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·
Two families of method
The first family measures how predictable the text is to a language model. Perplexity is the model's surprise at each word; burstiness is how much that surprise varies across sentences. GPTZero's original 2023 model worked this way, and its own explainer said a perplexity above about 85 was more likely human. The same page now carries an update: as of autumn 2023 GPTZero no longer uses perplexity and burstiness, having moved to a deep learning classifier. A more refined statistical approach, DetectGPT by Mitchell and colleagues (ICML 2023), looks at the curvature of the model's probability function rather than raw perplexity, and needs no training data at all.
The second family trains a classifier on large sets of labelled human and AI text and lets it learn whatever patterns separate them. Turnitin's FAQ states plainly that its model is not programmed to evaluate burstiness or perplexity and that individual predictions may not be explainable feature by feature. GPTZero, Grammarly and QuillBot all describe their current detectors as trained models of this kind, and QuillBot says its model still weighs predictability, sentence variation and repetitiveness among its signals.
Training data sets the baseline
A classifier only knows what it was shown. Grammarly's page says its model was trained on tens of thousands of texts. GPTZero says millions of documents across genres. ZeroGPT describes web, educational and proprietary synthetic datasets. If a tool saw little formal academic prose, or few non-native writers, it will score those inputs differently from a tool that did. The bias thread shows how far that can go.
Thresholds decide what a number becomes
Every tool ends with a probability and then has to turn it into a label. Where it draws the line is a business decision, not a fact about the text. QuillBot says that when a result is unclear its model leans towards "human" to reduce false positives. GPTZero says its output mapping is deliberately biased towards false negatives over false positives and that its "high confidence" band is tuned to under 1% error. Two tools with identical raw scores can therefore show different verdicts.
What each tool's score actually represents
Tool Method (as described by the vendor) What the score means What a "70%" means here GPTZero Sentence-by-sentence deep learning classifier; earlier versions used perplexity and burstiness Document class (human, mixed, AI) with class probabilities and a confidence band 70% probability of the AI class, read alongside the confidence label, not a share of the text ZeroGPT Multi-stage deep learning model the vendor calls DeepAnalyse; claims 98.4% accuracy and under 1% false positives Percentage of the text flagged as AI, with highlighted sentences About 70% of the text was flagged, judged sentence by sentence QuillBot Trained language-model classifier weighing predictability, sentence variation and repetition; 80-word minimum 0 to 100% likelihood the text was AI-generated, plus sentence highlights and AI-refined labels 70% likelihood, with uncertain cases pushed towards human Grammarly Segment-by-segment trained classifier Percentage of the text that "appears to be AI-generated" About 70% of segments matched AI patterns What that means in practice
- Compare labels, not numbers. A 70% from GPTZero and a 70% from ZeroGPT are different units. Only the tool's own documentation tells you which.
- Short text inflates disagreement. QuillBot requires 80 words and Turnitin 300; below that, a single sentence can swing the score.
- Benchmarks exist but vendors pick them. The RAID benchmark (Dugan and colleagues, ACL 2024) found detectors were easily fooled by sampling changes and adversarial edits, which is why a watermark or a writing record is stronger evidence than any score.
1 more reply
Most helpful first- Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·
If you want to see the threshold effect for yourself, run one long human-written paragraph through two tools, then delete the most formulaic sentence and run it again. The percentage tools (ZeroGPT, Grammarly) tend to move in small steps because one sentence is a small share of the text. The probability tools (GPTZero, QuillBot) can jump from one label to another when a borderline document crosses the cut-off. Neither behaviour is wrong; it just shows that the number is downstream of a decision the vendor made, not a property of your writing.
Write something first.
Give people something to work with: at least 30 words on what happened and what you tried.
That is too long. Keep it under 6,000 characters.
Write the question as the title, 15 to 140 characters, no links.
Pick a category.
Add a name (2 to 40 characters, no links).
That email address does not look right.
That was quick. Read the thread, then try again.
The form expired. Reload the page and post again.
Something went wrong with the form. Reload and try again.
Please complete the check and post again.
Limit reached for now. Try again later.
This thread is closed to new replies.
This thread no longer accepts replies.
Something went wrong with the form. Reload and try again.
Post a reply or question first (name and email), then this browser can vote, edit and accept answers.
You cannot vote on your own post.
Only the author (within 30 days) or the forum team can do that.
That email belongs to a forum team account. Use your sign-in link instead.
New accounts are paused for the moment. Try again later.
That was already posted.