# Do AI humanizers and paraphrasers actually beat AI detectors?

Source: https://forum.global100.org/q/do-ai-humanizers-and-paraphrasers-actually-beat-ai-detectors/
Site: Global 100 Forum, category Humanizers and rewriting
Published: 2026-09-15. Updated: 2026-09-21. Replies: 1.

## Question by Maya Lindqvist (staff), 2026-09-15

Often, yes, at least for a while: published attacks show that paraphrasing AI text sharply cuts detection rates on most style-based detectors. The caveat is that vendors now train specifically against paraphrasers and bypassers, the rewrite can damage meaning, and a tool that beats one detector today may be caught by the same detector next month.

## What the research shows

Two 2023 papers set out to stress-test detectors. [Krishna and colleagues](https://arxiv.org/abs/2303.13408), at NeurIPS 2023, built an 11 billion parameter paraphraser called DIPPER that rewrites whole paragraphs while controlling how much wording and ordering change. Run over text from three language models, it evaded watermarking, GPTZero, DetectGPT and OpenAI's own classifier. The headline figure: DetectGPT's accuracy fell from 70.3% to 4.6% at a fixed 1% false positive rate, without appreciably changing the meaning.

[Sadasivan and colleagues](https://arxiv.org/abs/2303.11156), in "Can AI-Generated Text be Reliably Detected?", later published in Transactions on Machine Learning Research, used a recursive attack: paraphrase the output, then paraphrase the paraphrase. On passages of roughly 300 tokens it cut detection across watermark-based, neural, zero-shot and retrieval-based detectors, with text quality degrading only slightly in many cases. The same paper shows spoofing, where human text is made to look AI-generated to a watermark detector.

The [RAID benchmark](https://arxiv.org/abs/2405.07940) (Dugan and colleagues, ACL 2024) widened the test to over 6 million generations across 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies, run against 8 open and 4 closed-source detectors. Its conclusion is blunt: current detectors are easily fooled by adversarial attacks, changes in sampling, repetition penalties and unseen generative models. A humanizer is, in effect, an unseen model plus an adversarial attack, applied at once. For the baseline these attacks erode, see [the accuracy thread](/q/how-accurate-are-ai-text-detectors-really/).

## What the vendors say

Turnitin's [AI writing detection FAQ](https://guides.turnitin.com/hc/en-us/articles/28477544839821-Turnitin-s-AI-writing-detection-capabilities-FAQs) states that its model can identify text that was AI-generated and then modified by "AI paraphraser or bypasser (also called humanizers) tools", that this runs automatically on every English submission where the feature is enabled, and that it will not name the tools it trained against because a list would help students evade it. Grammar-only changes from tools like Grammarly are generally not flagged; Grammarly's generative paraphrasing and rewriting features probably will be.

Two other lines deserve attention. Turnitin keeps its document-level false positive rate under 1% by accepting that it will miss some AI text, and it lists "text that has been paraphrased without developing new ideas" among the patterns that can produce false positives. Heavy paraphrasing can push a detector either way.

## What humanizers actually change

Style-based detectors measure how predictable each next word is and how much sentence length and structure vary. A paraphraser attacks exactly those signals: less probable synonyms, reordered clauses, split or merged sentences, broken-up repetition. Watermarks erode because the planted word choices get replaced, which is why [the watermarking thread](/q/does-watermarking-ai-generated-text-actually-work/) lists paraphrasing as that approach's main weakness too.

**Four ways people rewrite AI text and what each one costs**

| Technique | Effect on detectors | Effect on meaning and quality | Risk |
| --- | --- | --- | --- |
| Light manual editing | Small; unedited stretches still score as AI | Preserved; reads naturally | Low if disclosed, rarely changes a verdict |
| Single-pass AI paraphrase | Large drop on older or zero-shot detectors (the DIPPER result) | Mostly preserved; some odd synonyms | Vendors now train on it; can be flagged as paraphrased AI |
| Recursive paraphrasing | Largest drop in published attacks, including on watermarks | Quality falls each pass; nuance and citations drift | Less coherent, easier for a reader to spot |
| Dedicated bypass tools | Works until the detector is retrained | Varies; output can read awkwardly | Treated by institutions as evidence of intent |

## What that means in practice

- **Evasion is real but temporary.** The research shows paraphrasing beats most detectors as tested; the vendor statements show the target moves.
- **The cost lands on the text.** Recursive rewriting trades meaning and readability for a lower score.
- **Paraphrasing cuts both ways.** A heavily paraphrased human draft can itself trip a detector.
- **Nothing here produces proof.** A low score after rewriting says nothing about who wrote the ideas.

## Reply 1 by Maya Lindqvist (staff), 2026-09-21

One practical point that gets lost in the cat-and-mouse framing: if you are rewriting your own genuine draft and a paraphraser is the only way to get a "human" score, the problem is the detector, not your writing, and the fix is evidence rather than more rewriting. Keep the version history, the notes and the sources, because those survive any change in detector models. If, on the other hand, the text started as AI output, a paraphraser does not change what it is, and institutions increasingly treat bypass tools as an aggravating factor rather than a grey area.

---
Cite as: Global 100 Forum, "Do AI humanizers and paraphrasers actually beat AI detectors?", https://forum.global100.org/q/do-ai-humanizers-and-paraphrasers-actually-beat-ai-detectors/, accessed 2026-10-11.
