Invisible Ink
What happens when something you created passes through an AI system? Does it become permanently marked as "AI-generated"? Can a human-written document acquire an AI watermark simply because Claude edited it? And what happens when content moves between Claude, ChatGPT, Microsoft Word, Photoshop and PDF?
Those questions are becoming considerably more important.
In August 2026, Anthropic announced that future Claude models will generate text containing a digital watermark — an invisible statistical signal designed to help determine whether Claude was involved in producing a piece of writing.
Anthropic is also implementing provenance credentials for certain files produced or processed through Claude.
At first glance, this sounds straightforward:
The reality is considerably more nuanced.
Text and images are not being marked in the same way. A Claude watermark is not necessarily a permanent digital tattoo attached to everything that passes through the system. Moving content between different AI models can change the situation substantially. And perhaps most importantly, detecting Claude's involvement does not establish that Claude was the original author of the ideas contained within a document.
Understanding those distinctions will matter enormously as AI becomes embedded in writing, publishing, journalism, education, marketing, design and ordinary office work.
Why Is Anthropic Doing This?
Anthropic says it is implementing text watermarking partly in response to emerging European regulation.
In July 2026, Anthropic and other AI providers signed the European Union's Code of Practice on Transparency of AI-Generated Content. Anthropic says the requirements include mechanisms for marking AI-generated text.
Rather than limiting the technology geographically, Anthropic says it intends to apply the watermark globally at launch.
The objective is understandable. As generative AI becomes increasingly capable, distinguishing human-created material from machine-generated material becomes increasingly difficult. Watermarking potentially provides another tool for establishing provenance.
But the word watermark can create the wrong mental picture. Claude isn't necessarily hiding a little digital label between the letters of your paragraph. For text, something much more interesting is happening.
Think about a traditional watermark in a banknote. You can spend the bill normally. The watermark doesn't interfere with reading the denomination or using the money. But someone who knows what to look for can examine the bill and find evidence about its origin.
AI watermarking attempts something conceptually similar. With Claude-generated text, however, there isn't necessarily a secret message hidden inside the file. Instead, the evidence exists in statistical patterns created by the words — more technically, tokens — the AI chose while writing.
That distinction becomes extremely important later.
How Claude's Text Watermark Works
Anthropic says its text watermark is based on a version of SynthID-Text, a watermarking technique developed by Google DeepMind and described in peer-reviewed research published in Nature.
Large language models don't ordinarily construct sentences by retrieving complete sentences from a database. They generate text incrementally. Given everything written so far, the model calculates probabilities for possible tokens that could reasonably come next.
Suppose, in an extremely simplified example, the model reaches:
"The storm moved rapidly across the…"
Possible next words might include:
- city
- valley
- coast
- region
- landscape
All could be perfectly reasonable. A text-watermarking system can subtly influence those choices. SynthID-Text modifies the model's sampling process so that its token selections collectively produce a detectable statistical pattern while still generating natural-looking language.
The reader doesn't see the pattern. The sentence still means what it means. But across enough generated text, specialized detection software can examine those token choices and determine whether their statistical distribution is consistent with watermarked generation.
This is why Anthropic's text watermark should not be imagined as "PROPERTY OF CLAUDE" hidden somewhere inside a Microsoft Word file. The watermark is associated with the statistical characteristics of the generated language itself.
Imagine two writers choosing among several equally acceptable words. One chooses normally. The other follows an invisible rule: whenever several words would work equally well, slightly favour certain choices according to a secret pattern.
After one sentence, you probably couldn't tell. After hundreds or thousands of words, however, someone who knows the secret pattern could examine the choices and say: "This sequence is unusually consistent with our pattern."
Nothing is written in invisible ink. The pattern exists in the choices themselves. That is much closer to how statistical text watermarking works.
What Exactly Does Detection Prove?
This is one of the most important distinctions in the entire discussion. A positive watermark result should not automatically be interpreted as "Claude authored this document."
A more defensible interpretation is: "The text contains statistical evidence consistent with Claude having generated some or all of this language."
Consider a novelist who spends six months developing a story, characters, dialogue, research and plot, then asks Claude: "Rewrite this paragraph to improve its readability."
Claude may generate new wording. That wording could contain Claude's watermark. But Claude did not necessarily originate the story, the characters, the research, the argument, the underlying ideas, the factual investigation, or the author's creative intent.
The watermark speaks to generation provenance, not necessarily intellectual authorship.
Anthropic itself makes an important related distinction: its credential indicates Claude's involvement. It does not identify a particular user, organization or conversation.
That distinction is likely to become increasingly important in education, journalism, publishing and intellectual-property disputes.
What If I Wrote the Material Myself and Only Used Claude to Edit It?
This is where the issue becomes particularly interesting.
Suppose you write a 5,000-word article entirely yourself. You then paste it into Claude and ask: "Correct the spelling and punctuation but don't change my wording."
Simply passing text through Claude does not mean Claude retroactively generated your original words. The statistical watermark is introduced through Claude's generation process.
Now change the instruction: "Rewrite this article to make it more professional and concise."
Claude is now generating substantial new language. Those generated passages may contain the statistical watermark. The intellectual origin of the article may still be entirely yours, while portions of its linguistic expression have been generated by Claude.
Imagine you write a newspaper column and send it to a human editor. The editor changes "The committee really didn't understand what residents wanted" into "The committee appears to have misunderstood residents' concerns."
The editor created that particular sentence. But we wouldn't normally say the editor therefore conceived and authored the entire article.
AI introduces a technological version of this old distinction — except now we may have tools capable of detecting some of the editor's linguistic fingerprints.
What Happens When Content Moves From ChatGPT to Claude?
Consider this workflow:
Suppose ChatGPT generates an article. You paste it into Claude and ask Claude only to organize the existing material into sections.
If Claude leaves most of the original language unchanged, much of that language was not generated by Claude in the first place. But if Claude substantially rewrites the article, Claude is producing a new sequence of tokens. That newly generated language can contain Claude's watermark.
The fact that ChatGPT created an earlier draft doesn't prevent Claude-generated portions of the later draft from carrying Claude's statistical signal. Again, the detector isn't reconstructing the philosophical history of the document. It is examining the resulting language for statistical evidence.
Now Reverse the Process
You write the original article. Claude substantially edits it. You then give Claude's version to ChatGPT and instruct ChatGPT to completely rewrite the article.
What happens to Claude's watermark? This exposes one of the inherent limitations of statistical text watermarking.
Anthropic explicitly acknowledges that editing can affect its watermark. According to the company, light editing probably will not completely remove it, while a complete rewrite in which the words are replaced will.
That makes sense given how the technology works. If ChatGPT generates entirely new wording, it is creating a different sequence of tokens using a different model and generation process. There is no little Anthropic tag being physically copied from sentence to sentence.
Therefore, substantial regeneration by another model can disrupt or eliminate the statistical pattern that Claude's detector is looking for. That isn't necessarily a security vulnerability. It is an inherent consequence of what the watermark actually is.
Imagine Claude deals 500 playing cards according to a subtle pattern. An expert who knows the pattern can examine the sequence afterward and recognize it.
Now someone changes five cards. Much of the pattern remains. Change 100 cards and the evidence becomes weaker. Shuffle and completely redeal the entire deck using another dealer, however, and Claude's original pattern may no longer exist.
Text watermarking works much more like this than like permanently stamping someone's name onto a document.
Copying Is Not the Same as Regenerating
This distinction deserves particular emphasis. Suppose Claude produces a watermarked paragraph.
| Action | Effect on the watermark |
|---|---|
| Copy and paste it into Word | Words remain essentially identical, so the statistical pattern remains. |
| Copy and paste it into a PDF | Changing the container doesn't inherently change the linguistic sequence. |
| Paste it into ChatGPT and ask a question about it | Merely transmitting the text through another service doesn't inherently rewrite the words. |
| Ask ChatGPT to substantially rewrite it | Another model generates a new token sequence — the statistical characteristics can change substantially. |
The important variable for text isn't therefore simply "What software did this document pass through?" It is: "Who or what generated the actual language we are examining?"
Short Text Creates Another Problem
Statistical detection needs evidence. A detector examining "Thank you very much. I'll see you tomorrow." has vastly less statistical information available than one examining a 3,000-word essay.
Google DeepMind specifically notes that SynthID-Text performs best with longer, more diverse model-generated responses. That makes intuitive sense — the fewer token choices available for analysis, the less statistical evidence exists from which to draw conclusions.
Consequently, watermark detection should not be treated like DNA identification. It is probabilistic evidence. That distinction matters.
Suppose you suspect someone is using a weighted coin. They flip it three times: heads, heads, heads. Suspicious? Perhaps. Proof? Certainly not.
Now suppose they flip it 10,000 times and get heads 8,700 times. You have much stronger statistical evidence.
AI text watermark detection faces a similar problem. Longer passages provide more observations from which the detector can evaluate a pattern.
Images and Other Files Are Different
This is perhaps the greatest source of confusion surrounding Anthropic's announcement. Anthropic isn't handling supported files in precisely the same way as generated text.
When Claude produces certain supported file types, including formats such as JPG and SVG, Anthropic says it will attach Content Credentials using the C2PA standard — the Coalition for Content Provenance and Authenticity.
Rather than subtly altering linguistic probabilities, C2PA establishes a cryptographically verifiable provenance record associated with a digital asset. A C2PA manifest can contain assertions describing an asset's provenance and editing history. Those claims can be digitally signed and cryptographically bound to the asset.
Anthropic describes its credential as indicating that a supported file was made or processed with Claude. Notice the wording. Not necessarily "Claude created everything you see here." Instead: "Claude participated in this asset's production or processing." That is a much more precise claim.
Claude's text watermark is somewhat like recognizing a characteristic pattern in someone's handwriting. A C2PA Content Credential is more like a passport travelling with a digital asset — it can contain authenticated statements about where the asset came from and what happened to it.
These are fundamentally different technologies. Text watermark: evidence encoded statistically through generation choices. Content Credential: provenance information cryptographically associated with a digital asset.
What Happens When an Image Is Edited?
This depends heavily upon the software and workflow. C2PA was specifically designed to support provenance across creation, editing and publication workflows.
A compatible application can read existing provenance information and potentially add subsequent information describing additional processing. This can produce something closer to a chain of custody. For example, conceptually:
could produce a provenance history documenting stages in the asset's life. Likewise, an AI workflow could potentially record that an asset was processed using an AI system.
However, Content Credentials are not magic. Applications and platforms must actually support and preserve the relevant provenance data.
The C2PA specification itself distinguishes between the content and assertions made about its provenance. Cryptographic mechanisms help establish whether claims are properly associated with the asset and whether they have been tampered with.
Importantly, C2PA's own guiding principles warn against treating provenance information as a judgment about whether content is inherently trustworthy. A valid provenance record tells us something about where content came from and what happened to it. It does not automatically tell us whether the content is true.
A photograph can have impeccable provenance and depict something misleading. An AI-generated illustration can have impeccable provenance and be completely harmless. A human-created image can have no provenance metadata whatsoever and still be entirely authentic.
Provenance and truth are related questions, but they are not the same question.
What About Word Documents and PDFs?
Here we need to distinguish between two separate things.
The Words Inside the Document
If Claude generated substantial portions of the text, the statistical characteristics of that language can potentially survive being copied into Word or converted into PDF because the underlying words remain substantially unchanged. The watermark doesn't require the .docx extension to survive — it resides in the statistical characteristics of the language.
The File Itself
C2PA provenance is a different layer. Whether a Content Credential survives conversion, export, editing or publication depends upon the particular file format, software and workflow involved.
Therefore: DOCX/PDF conversion does not inherently erase a linguistic watermark. But file provenance metadata and cryptographic credentials are a separate technical question. Conflating those two mechanisms leads to considerable misunderstanding.
Can Anthropic Identify Who Used Claude?
Anthropic says its watermark does not contain identifying information about the person, organization or Claude conversation responsible for the text. That is an important privacy distinction.
A detection system could potentially conclude: "This passage is statistically consistent with Claude-generated text." That does not inherently mean it can conclude: "John Smith generated this at 3:42 p.m. on August 12 using account X." Those are entirely different capabilities.
Anthropic has also announced that it intends to provide a watermark-detection API. The precise implementation details and public availability of that detector will therefore deserve scrutiny as the system is deployed.
Is This the Same as Existing "AI Detectors"?
No. That distinction is extremely important.
Many conventional AI-writing detectors attempt to infer whether text looks like AI-generated writing. They may examine characteristics such as predictability, linguistic patterns, sentence variation or other statistical properties.
A watermark detector instead looks for a pattern intentionally introduced during generation. Those are conceptually different approaches. One asks: "Does this writing resemble AI output?" The other asks something closer to: "Does this writing contain the statistical signature our generation system intentionally produces?"
That potentially makes watermark-based detection much more meaningful. But it still does not transform probabilistic evidence into absolute proof.
NIST's work on synthetic-content transparency similarly recognizes that watermark systems have competing requirements including detectability, robustness, security and low distortion.
Could Human Writing Accidentally Be Identified?
This is where detector design becomes crucial. Any statistical detection system must contend with false positives and false negatives.
A false positive occurs when human-written text is incorrectly identified as watermarked. A false negative occurs when genuinely watermarked text isn't detected.
Responsible deployment therefore requires thresholds, confidence levels and careful interpretation. A detector returning evidence consistent with a watermark should not automatically become an accusation of plagiarism, academic misconduct or deception. Context matters.
The consequences of getting that distinction wrong could be substantial.
Imagine an airport scanner flags a suitcase. The alert means: "Something here matches characteristics we were instructed to detect." It does not necessarily mean: "We have conclusively proven that this passenger committed a crime."
Detection should trigger interpretation and further investigation — not automatically become a verdict. Statistical AI watermarking deserves the same caution.
The Authorship Problem
Watermarking technology may ultimately force society to confront a question that predates artificial intelligence:
Consider five people.
Calling all five examples simply "AI-generated content" destroys important information. They represent radically different levels of human intellectual and creative involvement.
Watermarking can potentially tell us something about generation. It cannot, by itself, resolve questions of authorship, originality, ownership, intellectual contribution or creative responsibility. Those questions require context.
This Matters Far Beyond School Essays
The consequences extend into virtually every knowledge industry.
Journalism
A journalist may conduct interviews, obtain documents, investigate a story and construct the argument, while using AI to improve readability. Detection of AI-generated sentences doesn't invalidate the journalism.
Business
An executive may write a proposal and use AI to reorganize it. A watermark does not mean the business strategy originated with Claude.
Academia
A researcher may use AI for permitted editorial assistance. Universities will need policies capable of distinguishing unauthorized generation from legitimate language assistance.
Publishing
Authors increasingly use grammar checkers, developmental editors, transcription systems and AI-assisted writing tools. Determining where editing ends and authorship begins will become increasingly complicated.
Graphic Design and Photography
Content Credentials may provide useful evidence of an asset's production history. But "AI was involved" can encompass everything from generating an entire image to performing a comparatively mundane processing operation.
Law
Courts, investigators and attorneys will eventually confront provenance evidence. They will need to understand what such evidence actually establishes — and, equally importantly, what it does not.
A Watermark Is Evidence, Not a Verdict
This may be the single most important principle for the public to understand. A watermark should be treated as one piece of provenance evidence. It does not inherently establish:
- plagiarism
- deception
- copyright ownership
- factual accuracy
- intellectual authorship
- academic misconduct
- the identity of the user
- the percentage of human creative contribution
C2PA's own principles make a related point: provenance standards are intended to provide verifiable information about content history, not declare whether the underlying material is "good" or "bad." That is an exceptionally important distinction.
The Bigger Picture: From Detection Toward Provenance
For several years, public discussion has revolved around the question: "Can we detect AI?" That may ultimately be the wrong question.
The more useful question could become: "Can we understand the provenance of this content?"
There is an important difference. Detection creates a binary cultural instinct — human or AI? Provenance recognizes the reality:
Modern creative production is increasingly a chain. The same has actually been true for decades. A professional photograph may pass through a camera, RAW processor, Photoshop, colour-grading software, layout application and publishing platform. An article may pass through a journalist, editor, copy editor, spellchecker, legal department and content-management system.
AI simply adds another participant to those chains. A mature provenance system should therefore tell us what happened, rather than attempting to reduce a complex creative history to a simplistic label.
The Paradox of AI Watermarking
There is an interesting paradox here. The better AI becomes integrated into everyday software, the less meaningful the phrase "AI-generated" may become without additional context.
Microsoft Office, Adobe software, search engines, smartphones, cameras, operating systems, email clients and creative applications increasingly contain AI functionality.
If a photographer removes three pedestrians from the background of a photograph using generative fill, is the photograph AI-generated? If an author asks Claude to restructure three paragraphs in a 70,000-word book, is the book AI-generated? If an accountant uses AI to summarize a spreadsheet, is the financial analysis AI-generated? If a journalist asks an AI system to correct punctuation, is the article AI-generated?
Technically, AI participated. But participation and authorship are not synonymous.
This is why provenance systems may ultimately prove more useful than binary labels.
What Consumers Should Remember
When you encounter claims about AI watermarking, keep several principles in mind.
- Claude's text watermark is statistical. It exists through patterns in generated token choices rather than simply being a hidden label attached to a Word document.
- Simply passing existing human-written text through Claude does not retroactively make Claude the author of those words.
- When Claude substantially rewrites material, the newly generated language may contain its watermark.
- Light editing may leave enough of the statistical signal to remain detectable, while extensive rewriting can weaken or eliminate it.
- Copying Claude-generated text into Word or PDF doesn't inherently remove the statistical characteristics of the language.
- Supported images and other files can use a fundamentally different mechanism — C2PA Content Credentials — that records provenance through cryptographically verifiable metadata.
- Provenance does not equal truth.
- AI involvement does not necessarily equal AI authorship.
- A detector should provide evidence — not judgment.
Where This Is Heading
Anthropic's announcement represents something larger than a technical modification to Claude. It is part of an emerging infrastructure for digital provenance.
Google is developing SynthID. Anthropic is adopting statistical text watermarking and C2PA Content Credentials. Camera manufacturers, software developers, publishers and technology companies are increasingly experimenting with C2PA-compatible provenance systems. Governments are beginning to regulate transparency around synthetic media.
The direction of travel is becoming clear. We are gradually constructing something resembling a chain of custody for digital information.
Done well, that could be enormously valuable. It could help journalists authenticate photographs. It could help consumers identify synthetic media. It could help creators document their workflows. It could make sophisticated misinformation more difficult to disguise. And it could provide transparency when artificial intelligence participates in creating content.
But it must be interpreted carefully. A society that treats every AI provenance marker as evidence of dishonesty risks turning a transparency technology into an accusation machine.
The technologically sophisticated question isn't simply: "Was AI involved?"
Increasingly, the better questions will be: How was AI involved? What did the human contribute? What did the machine contribute? What happened to the content afterward? And what, precisely, does the available evidence allow us to conclude?
Those questions acknowledge what modern digital creation increasingly is: not purely human, not purely machine, but a chain of tools, decisions, edits, ideas and authorship whose history is becoming technologically possible to document.
And understanding that history — rather than merely attaching an "AI-generated" label to the final product — may prove to be the real value of digital provenance.


