Claude Now Signs Everything It Writes: The Invisible Watermark That Survives Copy-Paste
You ask Claude for some text, copy it, paste it into a PR, a client email, a blog post. Until last week that was anonymous text. Not anymore.
Since August 2, 2026, every new Claude model ships with an invisible watermark embedded in the text itself. It's not metadata, not a header, not a weird character: it's a statistical signal stitched into the token sequence, one that rides along through copy-paste and, according to Anthropic itself, "may persist through some editing."
The detail almost nobody talked about: this wasn't announced in a blog post. It showed up in a help center article, How Claude marks AI-generated content. In this post I explain what it actually is, where it applies (spoiler: the API and Claude Code too), what detection proves and what it doesn't, and what actually removes the mark.
TL;DR
- What it is: an imperceptible watermark embedded in text generated by Claude models, plus signed provenance metadata (the C2PA standard) on supported files.
- Since when: models released on or after August 2, 2026 already ship marked; earlier models are getting support during a transition period.
- Where it applies: the model level. In other words: Claude Platform (API), Claude, Claude Code, Claude Cowork and Claude Tag, plus the partner clouds (AWS, Google Cloud, Microsoft Foundry).
- EU only? No. The requirement is European, the rollout is global.
- Can you turn it off? There's no documented API parameter, enterprise flag or Claude Code config for that.
- Official source: support.claude.com.
- Living post: I'll update this as Anthropic publishes the detector and expands coverage.
The context: why now, and why in a support article
The trigger has a name and a date. Article 50(2) of the EU AI Act requires generative AI providers to mark synthetic content in a machine-readable format, and the Article 50 transparency rules took effect on August 2, 2026. In June the European Commission finalized the Code of Practice on Transparency of AI-Generated Content, which turns the article into practice: watermark, metadata, a standardized icon and a free detection tool. Anthropic signed it.
It could have done the minimum: mark only what goes out to Europe, geolocate, add a region flag. It didn't. It applied it at the model level, for everyone, worldwide. From an engineering standpoint that makes total sense — segregating sampling behavior by jurisdiction is a caching, reproducibility and auditing nightmare. From the standpoint of someone using the API in Brazil, it means a European rule just became default behavior in your stack.
And the announcement? A help center article. No blog post, no thread, no keynote. That's not a conspiracy, it's an operational lesson: a change that alters the output of the models you have in production doesn't always arrive through the channel you watch. If your vendor radar is just the official blog, your radar has a hole in it. This kind of changelog reading and platform-contract change is weekly conversation over at Beer And Code, where the folks running AI in production share what broke before it turns into an incident.
It's not a zero-width invisible character
First, kill the myth. Half the timeline's initial reaction was "oh, it's that zero-width space, you can strip it with a regex." It isn't.
An invisible character (U+200B, Cyrillic homoglyphs, a creative non-breaking space) dies in any Unicode normalization. That would be a watermark that doesn't survive a .strip(). Anthropic doesn't say publicly which scheme it uses, but the description matches the established line of research from the last three years: biasing token sampling.
Here's how it works. At each generation step the model has dozens of plausible continuations with similar probabilities. "The system failed", "the system broke", "the system went down" — all fine. The watermark uses a secret key to decide, pseudorandomly, which subset of the vocabulary gets favored at that step. You don't notice any difference in quality. But the whole sequence now carries a statistical skew that only someone holding the key can measure.
It's the idea from A Watermark for Large Language Models (Kirchenbauer et al.), taken to scale by Google's SynthID-Text, which became a Nature paper with the tournament sampling scheme running in Gemini for millions of users.
Three practical consequences fall straight out of this, and they explain the rest of the post:
- Detection is a statistical test, not a flag check. The detector returns confidence, not an honest boolean.
- It needs length. Few tokens, little signal. A short sentence is undetectable by construction.
- It survives small noise, not a rewrite. Swapping two words barely moves the accumulated score. Paraphrasing everything throws the accumulated score away.
Where it applies: API, Claude Code, and no switch to turn it off
This is the part that matters if you build product. The marking happens at the model level, not the product level. The support article is explicit: it covers "output from supported models everywhere you use Claude," including the API, Claude, Claude Code, Claude Cowork and Claude Tag, and it also reaches you through AWS, Google Cloud and Microsoft Foundry (with signed-metadata support varying by platform).
Translated into architecture decisions:
- If you have a SaaS that generates product descriptions, ticket summaries or blog posts by calling Anthropic's API, the text that goes out to your customer goes out marked. With your vendor's signature, not yours.
- If you sell white label and the contract promises "original content," reread the contract before your customer's legal team rereads it for you.
- As of today there is no documented API parameter, enterprise control or Claude Code config to turn it off. Treat it as fixed platform behavior.
What about the code Claude Code writes? Anthropic doesn't go into detail. But the physics of the scheme gives you a hint: a sampling watermark depends on entropy. A for loop, a getter, a Laravel migration — these are stretches with very few plausible continuations. Where the model has almost no choice, there's almost nowhere to hide a signal. Until proven otherwise, don't count on a reliable watermark in short or highly idiomatic code. Long prose is the terrain where this works.
Files are a different story: signed C2PA
The second mechanism is different and worth separating out. For supported file types (.png, .jpg, .svg), Claude attaches digitally signed provenance metadata, using the open C2PA standard. A signature makes tampering detectable: if someone messed with the provenance, you can tell.
The weakness is different too. Metadata doesn't survive a screenshot, a format conversion, a re-save in an editor or an upload to a social network that strips EXIF. In other words: the text watermark is fragile to rewriting and resistant to format changes; C2PA is resistant to rewriting and fragile to format changes. They're complementary on purpose.
You can inspect it today, with the official tool from the Content Authenticity Initiative:
c2patool imagem-gerada.png
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãWhat detection actually proves (and what it doesn't)
This is where I want you to pay attention, because it's where the public conversation has already started getting it wrong.
Anthropic itself sets the limit: the presence of the mark indicates that the content may have been processed by Claude. It doesn't confirm authorship — because Claude frequently processes human material. You paste in something you wrote and ask it to "revise the tone": it comes out marked. You translate someone else's document: it comes out marked.
And the absence of the mark proves nothing either. The text may have been heavily edited, be too short, have gone through a format conversion, or come from a model that predates the transition period.
Put the two ends together and you land on the uncomfortable conclusion: this is a provenance signal, not proof of authorship. It's good for platforms labeling content at scale, for data pipelines avoiding training on synthetic output, for compliance teams documenting origin. It's not good for failing a student, cutting a freelancer or eliminating a candidate. Anyone who uses it as proof is going to get it wrong, and get it wrong against people who wrote the text by hand — the false positive risk was exactly the most repeated theme in the Hacker News discussion.
The test in practice: what erases the mark
There's no public detector from Anthropic yet — the company said it will publish technical details and a tool, with no timeline. So nothing here is a measurement of mine with numbers; it's what the public evidence supports today. I marked the confidence for each row.
| Action | Expected effect on the mark | Confidence |
|---|---|---|
| Copy and paste | Keeps it | High (official statement) |
Save as .txt / .md / paste into Word |
Keeps it | High |
| Swap 2 or 3 words | Keeps it | Medium-high |
| Rewrite a few paragraphs by hand | Weakens it | Medium |
| Translate pt → en → pt | Degrades it a lot | Medium-high |
| Paraphrase everything with another LLM | Destroys it in practice | High |
| Short snippet (a sentence, a commit message) | Undetectable | High |
| Screenshot / rasterized PDF | Dies with the text (OCR rebuilds an approximation) | Medium |
| File: conversion, re-save, social upload | C2PA metadata is gone | High |
The basis for translation and paraphrasing isn't a guess: Google's SynthID documentation, which is the best-documented comparable scheme, states that detector confidence drops sharply when the text is thoroughly rewritten or translated into another language. Makes sense: a translator resamples the entire text. There's no original sequence left to measure.
The honest summary is this: the mark survives laziness, not work. Paste it straight in and you're marked. Actually rewrite it and you're not. Which, come to think of it, is more or less the right incentive.
If you want to be ready to measure this properly when the detector ships, start saving a baseline now — the raw response, before any human editing:
# baseline: salve o output original ANTES de qualquer edição
curl -s https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d "{\"model\":\"$MODEL\",\"max_tokens\":2048,\"messages\":[{\"role\":\"user\",\"content\":\"$PROMPT\"}]}" \
> "artifacts/$(date +%s)-raw.json"
With the baseline saved, the day the detector shows up you test variations against the original instead of trying to reconstruct what happened.
What to do if you put LLM output in a product
A short checklist, in the order I'd do it:
- Map where Claude output crosses your boundary and reaches an end customer, an email, a document or a public API. That's the list that matters.
- Record provenance on your side. Don't depend on your vendor's watermark to know what your own application generated.
- Review contracts and terms that promise "original" or "not AI-generated" content.
- Adjust your product's disclosure policy. Article 50 also covers visible labeling for the end user, not just machine marking.
- Don't build punishment on top of an AI detector. Not a third-party one, and not the official one when it ships.
- If you generate images or SVG, decide deliberately whether your optimization pipeline preserves or destroys the C2PA. Today you're probably destroying it without knowing.
Item 2 is the cheapest and the most ignored. In Laravel you can solve it with one table and three lines at the moment you persist the generation:
AiOutput::create([
'surface' => 'produto.descricao',
'model' => $response->model,
'response_id' => $response->id,
'sha256' => hash('sha256', $text),
'chars' => mb_strlen($text),
'user_id' => auth()->id(),
]);
Your own provenance, versioned, auditable, without depending on anyone publishing any detector. It's a compliance concept turning into two columns in the database — which is how compliance should reach devs. If that's your context, the 18-item EU AI Act + NIST RMF checklist goes deeper on the rest of the conversation.
Quick FAQ
If I use the API through Bedrock or Vertex, do I escape it? No. The marking is at the model, and the support article explicitly names AWS, Google Cloud and Microsoft Foundry. What varies by platform is support for signed metadata on files, not the text watermark.
Are older models already in production marked? Not necessarily. The rule applies right away to models released on or after August 2, 2026; for earlier ones, Anthropic said it is adding support during a transition period. When in doubt, assume it may be marked.
Does this change response quality? According to Anthropic, it doesn't change meaning, quality or readability. In published schemes of this kind, the impact on perplexity is small. If you have versioned evals, now is a good time to rerun them and see for yourself.
Is there any legitimate way out? None documented. If the watermark is a real blocker for your use case — and legitimate cases exist, like journalism with a protected source — the path is to talk to Anthropic or evaluate a self-hosted open model, not a removal hack.
Conclusion
It's not the end of cheating with AI. It's the beginning of a provenance infrastructure — imperfect, asymmetric and full of rough edges, but actually running, at the largest enterprise-focused AI lab.
What changes for you has already changed: if you call Claude in production, your output carries your vendor's signature, with no way to turn it off, because of a law that may not even apply to you. And Anthropic told you about it in a help center article.
The next step is the public detector. When it ships, the question stops being "was this generated by AI?" and becomes who has the right to run that check, on whose text, with what consequence. That conversation is a lot harder than the technical one. In the meantime, the usual work: adopting AI with process instead of faith.
Updates
- August 11, 2026 — initial publication, based on Anthropic's support article (August 10, 2026). Public detector not yet announced. I'll update this section as Anthropic releases the detection tool and details coverage for code and earlier models.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã