Skip to content
Global 100 Forum

Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?

Staff answerWritten by Maya Lindqvist · 1 reply · updated

Short answer

They disagree because they are answering different questions with different models, different training data and different cut-offs. A "70%" from one tool is a probability, from another a share of sentences, and from a third a share of text. The caveat: none of them measure the same thing, so averaging or comparing the numbers is meaningless.

CiteGlobal 100 Forum, "Why do GPTZero, ZeroGPT, QuillBot and Grammarly give different AI scores?", https://forum.global100.org/q/why-do-gptzero-zerogpt-quillbot-and-grammarly-give-different-ai-scores/, accessed 2026-10-11.
Written by Maya Lindqvist · Edits the AI text detection and academic integrity sections ·
  1. Edits the AI text detection and academic integrity sections ·

    Two families of method

    The first family measures how predictable the text is to a language model. Perplexity is the model's surprise at each word; burstiness is how much that surprise varies across sentences. GPTZero's original 2023 model worked this way, and its own explainer said a perplexity above about 85 was more likely human. The same page now carries an update: as of autumn 2023 GPTZero no longer uses perplexity and burstiness, having moved to a deep learning classifier. A more refined statistical approach, DetectGPT by Mitchell and colleagues (ICML 2023), looks at the curvature of the model's probability function rather than raw perplexity, and needs no training data at all.

    The second family trains a classifier on large sets of labelled human and AI text and lets it learn whatever patterns separate them. Turnitin's FAQ states plainly that its model is not programmed to evaluate burstiness or perplexity and that individual predictions may not be explainable feature by feature. GPTZero, Grammarly and QuillBot all describe their current detectors as trained models of this kind, and QuillBot says its model still weighs predictability, sentence variation and repetitiveness among its signals.

    Training data sets the baseline

    A classifier only knows what it was shown. Grammarly's page says its model was trained on tens of thousands of texts. GPTZero says millions of documents across genres. ZeroGPT describes web, educational and proprietary synthetic datasets. If a tool saw little formal academic prose, or few non-native writers, it will score those inputs differently from a tool that did. The bias thread shows how far that can go.

    Thresholds decide what a number becomes

    Every tool ends with a probability and then has to turn it into a label. Where it draws the line is a business decision, not a fact about the text. QuillBot says that when a result is unclear its model leans towards "human" to reduce false positives. GPTZero says its output mapping is deliberately biased towards false negatives over false positives and that its "high confidence" band is tuned to under 1% error. Two tools with identical raw scores can therefore show different verdicts.

    What each tool's score actually represents

    Tool Method (as described by the vendor) What the score means What a "70%" means here
    GPTZero Sentence-by-sentence deep learning classifier; earlier versions used perplexity and burstiness Document class (human, mixed, AI) with class probabilities and a confidence band 70% probability of the AI class, read alongside the confidence label, not a share of the text
    ZeroGPT Multi-stage deep learning model the vendor calls DeepAnalyse; claims 98.4% accuracy and under 1% false positives Percentage of the text flagged as AI, with highlighted sentences About 70% of the text was flagged, judged sentence by sentence
    QuillBot Trained language-model classifier weighing predictability, sentence variation and repetition; 80-word minimum 0 to 100% likelihood the text was AI-generated, plus sentence highlights and AI-refined labels 70% likelihood, with uncertain cases pushed towards human
    Grammarly Segment-by-segment trained classifier Percentage of the text that "appears to be AI-generated" About 70% of segments matched AI patterns

    What that means in practice

    • Compare labels, not numbers. A 70% from GPTZero and a 70% from ZeroGPT are different units. Only the tool's own documentation tells you which.
    • Short text inflates disagreement. QuillBot requires 80 words and Turnitin 300; below that, a single sentence can swing the score.
    • Benchmarks exist but vendors pick them. The RAID benchmark (Dugan and colleagues, ACL 2024) found detectors were easily fooled by sampling changes and adversarial edits, which is why a watermark or a writing record is stronger evidence than any score.
    0
    ReplyLink

1 more reply

Most helpful first
  1. Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·

    If you want to see the threshold effect for yourself, run one long human-written paragraph through two tools, then delete the most formulaic sentence and run it again. The percentage tools (ZeroGPT, Grammarly) tend to move in small steps because one sentence is a small share of the text. The probability tools (GPTZero, QuillBot) can jump from one label to another when a borderline document crosses the cut-off. Neither behaviour is wrong; it just shows that the number is downstream of a decision the vendor made, not a property of your writing.

Write a reply

Plain text or simple Markdown. Links are nofollow. Your email is never shown.