Every word in a sentence carries a probability, some statistical likelihood of appearing given everything that came before it. A language model normally picks words based on that probability alone. A watermarking system quietly tilts the odds instead, nudging certain word choices just enough to leave a pattern behind, one invisible to a reader but recoverable by a detector that knows what to look for.
Understanding that mechanism is the first step to understanding what this tool is actually doing when it processes a piece of writing.
What Gets Marked and How
Google’s SynthID, currently the most widely deployed text watermarking system, works by biasing token probabilities during generation. Instead of choosing the statistically most likely next word every time, the model is nudged toward a specific subset of equally reasonable options, chosen according to a pattern only the detector knows. Across a long enough passage, that consistent nudging creates a signature a detection tool can recognize using a scoring method that compares observed word choices against the expected pattern.
The nudging is deliberately subtle enough that it should never make a sentence read as awkward or unnatural on its own. Among equally plausible word choices at any given point, the system simply favors a specific subset more often than pure probability alone would, which is precise enough to be recoverable statistically while staying invisible to a reader.
Why some text carries a stronger signal than other text
Length and openness matter
Watermarking works best on longer, more open-ended writing, essays, stories, emails, anywhere a model has real freedom in word choice. Short or highly factual text gives the system far fewer opportunities to nudge probabilities without changing the actual answer, which is why a one line factual response carries a much weaker signal than a full paragraph of narrative writing.
Not every model uses the same approach, or any approach at all
Google’s Gemini uses SynthID. OpenAI has researched text watermarking but has not confirmed deploying it in ChatGPT, and Anthropic’s Claude has never released a text watermarking system at all. Whether this matters for a specific piece of writing depends entirely on which tool generated the starting draft.
How Cleanup Actually Interacts With the Pattern
A few structural changes tend to disrupt a watermark’s underlying pattern:
- Reordering sentence structure, since the pattern depends partly on sequence
- Substituting words at points where the original choice was statistically nudged
- Restructuring paragraphs rather than editing individual sentences in isolation
- Varying sentence length and rhythm, which shifts the token distribution broadly
Why partial edits often are not enough
A watermark is distributed across many tokens throughout a passage, not concentrated in a single sentence. Editing a handful of sentences while leaving the rest of a paragraph untouched can leave enough of the original pattern intact for a detector to still register a match, even though the piece looks substantially rewritten to a human reader. That gap between how a piece reads and what its underlying statistics still show is exactly the problem a dedicated cleanup tool is built to close.
What This Means for a Genuinely Edited Piece of Writing
None of this is about disguising anything. It is about making sure a piece of writing that has genuinely been reworked, restructured, and personalized does not still carry a leftover statistical signature from a much earlier stage of its life, one that no longer reflects how the final version actually reads.
The distinction between a signature and a judgment matters here. A watermark says something about how a specific passage of text was generated at a specific moment. It says nothing about how much a person changed afterward, how original the final argument is, or how much genuine thought went into the finished piece. Treating a leftover pattern as evidence of anything beyond its narrow technical meaning is a misreading of what the technology actually measures.
What this looks like when it goes wrong
A writer who spends hours reworking a piece into something genuinely their own, then gets an unexpected flag from a detector picking up a stale pattern from the very first draft, is dealing with a real mismatch between effort and outcome. That mismatch is exactly what a dedicated cleanup step exists to close, restoring the writing’s statistical profile to match what a reader would actually experience reading it.
The mechanics behind text watermarking are more specific and more limited than most people assume, a token level pattern that varies by model, by content length, and by how much editing happened afterward. Knowing how the pattern actually works is what makes it possible to judge whether a piece of writing still carries one, rather than guessing.
Anyone dealing with this at real scale, checking large volumes of text for leftover statistical patterns, will likely end up exploring more of what an account like this offers beyond just this one specific tool.
That wider set of tools sits under Phrasly AI.
What this means in practice
Short passages and heavily edited text are consistently the hardest cases for any watermark detector to call with confidence, which is worth keeping in mind before treating a detection result on either of those as definitive. The pattern needs room to show up, and both short and heavily rewritten text give it very little.
That limitation cuts both ways. A short piece of writing is unlikely to trigger a confident watermark match even if it started as AI output, and a heavily edited piece is unlikely to trigger one even if a watermark was originally present. Detection confidence tracks length and how much editing happened, not some fixed property of whether AI was involved at all.
View 0 comments