Skip to content
Global 100 Forum

Does watermarking AI-generated text actually work?

Staff answerWritten by Tomas Reyes · 1 reply · updated

Short answer

It works when the model provider builds it in and the text is not heavily edited, which makes it far more reliable than guessing from style. It does nothing for text from models that do not watermark, and paraphrasing or translation can weaken the signal. It is a provider-side tool, not something a teacher can apply afterwards.

CiteGlobal 100 Forum, "Does watermarking AI-generated text actually work?", https://forum.global100.org/q/does-watermarking-ai-generated-text-actually-work/, accessed 2026-10-11.
Written by Tomas Reyes · Edits the deepfakes and provenance sections ·
  1. Edits the deepfakes and provenance sections ·

    How it works

    A text watermark nudges the model's word choices in a pattern that is invisible to readers but statistically detectable by someone who knows the key. The influential early design is Kirchenbauer and colleagues' "A Watermark for Large Language Models", presented at ICML 2023, which splits the vocabulary into "green" and "red" lists at each step and slightly favours green words. A detector counts green words and runs a statistical test.

    Google DeepMind described a production system in Dathathri and colleagues' "Scalable watermarking for identifying large language model outputs", published in Nature in October 2024. It changes only the sampling step, does not need the original model to detect the watermark, and the authors report it preserved text quality in a large live test. The paper also states the key limitation plainly: watermarks can be circumvented by editing or paraphrasing.

    Three ways to tell where text came from

    Approach Who controls it Can give positive proof? Weak point
    Style-based detector Anyone, after the fact No, only a probability (see accuracy thread) False positives, bias, easy to evade by editing
    Text watermark The model provider, at generation time Yes, when the key holder checks Only covers watermarked models; paraphrasing erodes it
    Signed provenance (C2PA-style) The producer or tool, at creation time Yes, for the file's history (see C2PA thread) Optional; stripped by screenshots and re-encoding

    What that means in practice

    • Coverage is the real constraint. A watermark only exists if the provider chose to add it. Text from open models, or from providers that do not watermark, carries nothing to detect.
    • Detection needs the key. Generally only the provider, or partners it shares detection with, can check for its watermark.
    • Editing erodes it. Short passages and heavily rewritten text may not carry enough signal for a confident result. The thread on humanizers and paraphrasers covers what the research says about deliberate rewriting.
    0
    ReplyLink

1 more reply

Most helpful first
  1. Tomas ReyesStaffEdits the deepfakes and provenance sections ·

    A useful way to place watermarking next to the other approaches in this forum: style-based detectors guess from the text, watermarks check for a signal the provider planted, and C2PA-style provenance records a signed history. Only the last two can produce strong positive evidence, and both depend on the producer opting in. There is a separate thread on C2PA and Content Credentials if you want the media side of the same idea.

Write a reply

Plain text or simple Markdown. Links are nofollow. Your email is never shown.