Skip to content
Global 100 Forum

Are AI detectors biased against non-native English writers?

Staff answerWritten by Maya Lindqvist · 1 reply · updated

Short answer

Yes, there is good evidence that they can be. The best-known study found that popular detectors flagged more than half of essays written by non-native English speakers as AI-generated, while scoring essays by native-speaking students almost perfectly. If your students or colleagues write in a second language, treat a high score with extra caution.

CiteGlobal 100 Forum, "Are AI detectors biased against non-native English writers?", https://forum.global100.org/q/are-ai-detectors-biased-against-non-native-english-writers/, accessed 2026-10-11.
Written by Maya Lindqvist · Edits the AI text detection and academic integrity sections ·
  1. Edits the AI text detection and academic integrity sections ·

    The study is Liang, Yuksekgonul, Mao, Wu and Zou, "GPT detectors are biased against non-native English writers", published in Patterns in 2023. The Stanford team ran seven detectors on two sets of human writing: TOEFL essays by non-native speakers and essays by US eighth-grade students. The detectors misclassified over half of the TOEFL essays as AI-written. The US student essays were classified almost entirely correctly.

    Key figures from Liang et al., 2023

    Writing sample Author group Share wrongly flagged as AI
    TOEFL essays Non-native English speakers More than half, across seven detectors
    US eighth-grade essays Native English speakers Near zero
    TOEFL essays after vocabulary enrichment Non-native English speakers Fell substantially, showing the detectors react to word choice

    Why it happens

    Many detectors lean on perplexity, a measure of how predictable each next word is to a language model. Writers working in a second language often use a smaller vocabulary and more common sentence patterns. That makes their text more predictable, which is exactly what the detectors associate with machine output.

    The researchers tested this directly. When they used a model to enrich the vocabulary of the non-native essays, the false positive rate fell. When they simplified the vocabulary of native-speaker essays, the false positive rate rose. The tools were reacting to linguistic sophistication, not authorship.

    What to do with that

    • Do not use a detector score as the basis for an accusation, and be especially careful where the writer is working in a second language.
    • Ask for process evidence such as drafts, notes or version history, which do not depend on how polished the prose is.
    • Check your tool's documentation for any published testing on non-native writing. Many vendors do not publish it.
    • Read the wider accuracy picture in how accurate are AI text detectors, really before setting any policy on scores.
    0
    ReplyLink

1 more reply

Most helpful first
  1. Maya LindqvistStaffEdits the AI text detection and academic integrity sections ·

    Some institutions reached the same conclusion from the policy side. When Vanderbilt University disabled Turnitin's AI detector in August 2023, one of the reasons it gave was that detectors are more likely to label text by non-native English speakers as AI-written. That is a useful document to share with a committee that is deciding how much weight to give a score.

Write a reply

Plain text or simple Markdown. Links are nofollow. Your email is never shown.