AI News

Claude’s Watermark Can Signal Contact. It Cannot Prove Guilt.

Anthropic’s announcement of invisible watermarks in Claude-generated text contains two sentences that deserve to be read together.

First, finding a Claude mark is “not fully conclusive.” It may indicate only that the content “may have been processed by Claude.” Second, failing to find one “doesn’t mean the content wasn’t AI-generated or processed.” Those are not hostile interpretations. They are Anthropic’s own limitations.

So what socially important conclusion is this detector actually qualified to support?

Not that Claude wrote the text. Not that a person cheated. Not that the work is unoriginal, false, plagiarized, or even meaningfully machine-authored. A positive result may mean Claude translated, proofread, summarized, reformatted, or otherwise handled material whose substance came from a human. A negative result may mean the text was rewritten, translated, mixed with other material, too short to test reliably, generated by an unsupported model, or stripped of file metadata.

Anthropic has announced a contact trace dressed for the part of a lie detector.

That contact trace may be useful. It could help a newsroom investigate a suspicious document, a platform study coordinated influence activity, or a researcher measure the spread of machine-generated material. The narrow provenance signal is defensible. Allowing it to acquire a social meaning far broader than its technical one would not be.

A provenance signal is being invited to audition as a lie detector. Schools, employers, publishers, clients, and platforms will not always ask the restrained question, “Does this passage carry a supported Claude watermark?” They will ask the question they actually care about: “Did this person write it?” Anthropic’s mark cannot answer that. Yet once it appears in a dashboard with a confidence score and an official logo, the temptation to pretend otherwise will be enormous.

That is the danger at the center of Claude’s new marking regime. The detector may be technically correct and institutionally calamitous at the same time.

A global mark, with the technical details to follow

Anthropic updated its help-center article on August 10, 2026, after signing Section 1 of the European Union’s Code of Practice on Transparency of AI-Generated Content. The company says Claude models launched in the EU on or after August 2, 2026 will support machine-readable marking at launch. It is working to add support to earlier models during the applicable transition period.

For supported models, the scope is expansive. Anthropic says marking will operate across the Claude API, the consumer service, Claude Code, Claude Cowork, Claude Tag, and deployments through AWS, Google Cloud, and Microsoft Foundry. It will apply wherever Claude is offered, not merely in the European Union. Text will receive an embedded watermark at the model level. Supported files, including formats such as SVG, PNG, and JPG, may carry signed provenance metadata using the C2PA standard.

The model-level choice matters. This is not a label added by one Claude interface or a piece of metadata that disappears when a response is copied from a browser. Anthropic says the text mark travels with copied text and may survive some editing. Downstream developers using the API or cloud platforms inherit the marking behavior of supported models.

But the company has disclosed almost nothing about the implementation that would allow outsiders to evaluate it. As of August 11, Anthropic had not published the algorithm or enough implementation detail to identify its watermark family; Claude-specific false-positive and false-negative rates; results by language, passage length, or subject; evaluations for code, mathematics, citations, or constrained factual output; detector thresholds; abstention rules; independent red-team results; or an appeal process. Its article says detection documentation is forthcoming.

That sequence is backwards. Anthropic has announced the global reach of the mark before explaining how the mark works, how often it fails, what a detector will display, who may use it, or how a person can contest the inference somebody else draws from it.

The page itself has not quietly supplied those missing answers. Its substantive article body on August 11 matched the earliest archived capture from August 10. The technical receipts remain a promise.

A contact trace is not a byline

The deepest problem is semantic, not statistical. Even a detector with no technical errors could still be used to reach the wrong human conclusion.

Imagine a researcher who spends months designing a study and writing a paper, then asks Claude to correct grammar. Imagine a non-native English speaker who translates an original essay into English. Imagine a developer who writes an application and asks Claude to refactor one function. Imagine a lawyer who supplies an original memorandum and requests a different file format. Imagine a journalist who asks Claude to shorten an interview transcript without changing its substance.

In each case a valid Claude mark could accurately report tool contact. It would not measure intellectual contribution.

That distinction is familiar in every other creative technology. A Microsoft Word file does not make Microsoft the author. A Photoshop edit does not reveal who conceived a photograph. A compiler does not own the program it translates. But an invisible text watermark arrives in a context already poisoned by the phrase “AI detection,” where evidence of software use is routinely mistaken for evidence that a human did not do the work.

Anthropic’s own examples make this category error impossible to ignore. The company warns that a mark may remain when Claude proofreads, translates, summarizes, or converts a file. Those activities span an enormous range. A summary can involve substantial generation. A spelling correction may involve virtually none. “Processed by Claude” compresses both into a single technical relationship.

The EU’s July 2026 implementation guidelines recognize the difference. They treat grammar correction, spellchecking, translation, formatting, minor polishing, transcription, and some assistive transformations as examples that may fall outside the provider marking obligation when they do not materially change meaning, style, or intent. By contrast, summarization and substantial rewriting can require marking. The distinction is imperfect and case-specific, but it at least attempts to separate assistance from authorship. The Commission’s guidelines do not pretend that every encounter with a language model has the same provenance significance.

Anthropic appears to be choosing a broader rule: embedded watermarks will apply to all generated text from supported models, worldwide, while its limitations contemplate marked proofreading and translation. Perhaps its eventual implementation will respect the legal exceptions at inference time. The announcement does not say. Until it does, the reasonable reading is that Anthropic prefers blanket consistency over a messy judgment about degree of transformation.

Operationally, that preference is understandable. Socially, it is dangerous. When the mark cannot distinguish a generated argument from a corrected comma, the detector delegates that distinction to whoever receives the result. That person may be a trained forensic analyst. It may also be a hurried professor, an automated moderation vendor, a suspicious client, or a manager looking for an easy answer.

People who rely on language assistance have particular reason to demand clarity. The Commission guidelines explicitly discuss technologies that transform authentic human input so people with disabilities can communicate. Translation and linguistic polishing are also more central to some people’s participation than others’. This does not prove that Claude’s policy will produce disparate harm. It does establish a foreseeable question: will the people who need assistive tools most be asked most often to prove that their ideas remain their own?

Anthropic cannot resolve that question by saying the mark does not alter ownership. The foreseeable risk is that somebody else will use the company’s signal to discount the human behind the work.

What the signal can and cannot support

Signal What it can support What it cannot support
A supported Claude text watermark is detected Evidence that the tested text carries a watermark associated with a supported Claude system, subject to the detector’s published limits Who originated the ideas; how much Claude contributed; whether use was allowed; plagiarism; deception; policy violation; truth; ownership; or misconduct
No supported Claude watermark is detected Only that the detector did not find a supported mark in the submitted sample Human authorship; absence of Claude processing; absence of other AI generation; or authenticity
Valid C2PA provenance metadata is present A signed provenance assertion is correctly formed, associated with the asset, and has not been altered undetectably That the asset is truthful, original, complete, ethically made, or accurately interpreted
C2PA metadata is missing or invalid The expected credential is absent, unavailable, stripped, changed, or cannot be validated That the asset is fake, malicious, or generated without AI
A detector abstains or reports low confidence The evidence is insufficient for a reliable classification Permission to round uncertainty up into guilt

The law is narrower than Anthropic’s policy

The European regulation behind this announcement is more careful than either its boosters or detractors often admit.

Article 50(2) of the AI Act requires providers of generative systems to ensure that synthetic audio, image, video, and text outputs are machine-readable and detectable as artificially generated or manipulated. The technical solutions must be “effective, interoperable, robust and reliable” as far as technically feasible, taking account of the content type, implementation cost, limitations, and the recognized state of the art.

That obligation has express boundaries. It does not apply to the extent a system performs an “assistive function for standard editing” or does not substantially alter the user’s input or its semantics. The Commission guidelines say standard editing includes small changes for grammar, readability, quality, format, accessibility, and similar publication needs. They identify AI translation and minor linguistic polishing as examples that can benefit from the exception. Material changes to meaning, style, structure, or intent fall on the other side.

Article 50 also assigns different duties to different actors. Providers such as Anthropic must implement machine-readable marking for covered generation. Deployers publishing AI-generated or manipulated text about matters of public interest have a disclosure obligation, but that obligation includes an exception when the publication has undergone human review or editorial control and a person or organization assumes editorial responsibility.

This division is important. The law does not say that every use of a generative tool destroys human authorship. It distinguishes technical marking at the provider layer from accountable publication at the deployer layer. It also recognizes that editing and human review change the meaning of AI involvement.

Article 50(2) is binding; the Code of Practice remains voluntary. The July 2026 Digital Omnibus on AI gave systems placed on the market before August 2 a four-month transition, until December 2, 2026, to meet the marking obligation. It also clarified the Code’s limited legal effect: the Commission may assess whether adherence is adequate, but the Code does not create a presumption of conformity. Signing can still reduce uncertainty. It does not transform every measure in the Code into a universal law of authorship or conclusively prove compliance.

Anthropic’s decision to deploy model-level marking worldwide is therefore a corporate policy choice as well as a compliance measure. The Brussels effect is familiar: a large regulated market establishes a rule, and global firms standardize around it rather than operate multiple systems. Sometimes this improves products everywhere. Sometimes it converts one jurisdiction’s compromise into everyone’s default without a separate argument about proportionality.

Anthropic has offered no such argument. Why should an editing exception recognized by EU law apparently disappear in a global implementation? Why should developers in Canada, writers in India, or businesses in Japan receive marked output because a model was launched into the European market? Why should a customer outside the EU have no stated choice?

“Consistency” is not an answer. Consistency is an engineering property. Proportionality is a governance judgment.

Watermarking works. Adversaries work around it.

The lazy critique says text watermarks are fake technology. That is wrong.

Google DeepMind’s peer-reviewed SynthID-Text work demonstrated that a generative watermark can be deployed at enormous scale. The method modifies token sampling rather than adding visible labels or necessarily inserting odd Unicode characters. In a live experiment covering nearly 20 million Gemini responses, Google reported no statistically significant change in user feedback between watermarked and unwatermarked outputs. Controlled tests also found no significant preference difference across grammaticality, relevance, correctness, helpfulness, and overall quality for the non-distortionary configuration.

That is serious evidence. A provider-applied watermark also has an important advantage over generic AI-text classifiers. Instead of guessing authorship from prose style, it tests for a signal deliberately introduced during generation. Under nominal conditions, especially with sufficiently long and varied text, that can be meaningful forensic evidence.

But meaningful is not magical. Google’s own plain-language explanation says SynthID works best on longer, diverse responses. Confidence can fall sharply after thorough rewriting or translation, and performance is weaker on factual or highly constrained prompts because the model has fewer acceptable token choices. The Nature paper identifies stealing, spoofing, scrubbing, and edits as continuing limitations. It also describes a real trade-off: configurations that push harder for detectability can sacrifice quality, while stronger non-distortion guarantees can reduce detectability or diversity.

Other research sharpens the problem. At EMNLP 2024, Saksham Rastogi and Danish Pruthi showed that with limited black-box access to generations, researchers could reverse-engineer aspects of watermark behavior and make paraphrasing attacks substantially more effective. A 2025 paper in Transactions on Machine Learning Research tested recursive paraphrasing across watermark, classifier, and retrieval approaches and found that detection could degrade under attack.

Then there is spoofing. An EACL 2026 paper introduced DITTO, a framework in which a smaller model learned and imitated the watermark signal of a studied model through knowledge distillation, creating authentic-looking marks associated with the target. The finding does not establish that Claude’s undisclosed watermark can be copied, but it makes imitation a credible threat model rather than a science-fiction objection.

These findings define watermarking’s proper scope. A watermark can be one clue in a layered investigation. It is poorly suited to become an automatic verdict in an adversarial environment.

The asymmetry matters more than any single benchmark. A motivated operator spreading disinformation has an incentive to paraphrase, translate, mix, regenerate, test, and adapt. A student who uses Claude for legitimate grammar correction, or a developer who pastes a generated helper function into a larger codebase, may do none of those things. The bad actor attacks the mark. The ordinary user preserves it.

A control that is most durable among the compliant and most contested among the sophisticated risks inverting accountability. It catches tool contact more readily than intent.

A private detector with public consequences

Anthropic says it will help users and third parties detect Claude’s marks. It has not yet said what “help” means.

The final EU Code of Practice exposes how unsettled this area remains. It says no single marking technique will suffice in many cases and anticipates combinations of metadata, watermarks, and supplementary methods. It defines text shorter than 200 tokens as “very short” under the present state of the art. Free-form text longer than 200 tokens should still be watermarked, but the Code acknowledges lower reliability than for very long text and allows access to the relevant detector to be restricted to verified experts when results could mislead the public.

The compliance framework itself contemplates an invisible mark applied broadly while access to interpreting that mark may be limited because the interpretation is not yet reliable enough for public use.

The Code also contains safeguards that should become the minimum public specification for Anthropic’s implementation. Detection results should identify whether they come from metadata, a watermark, a forensic method, or another technique. They should be comprehensible. A user should be able to download a digitally signed result containing a hash of the tested content, an identifier for the detection service, and a timestamp. Services that require uploads should use data minimization, protect confidentiality, process material only for detection, and delete it immediately afterward, subject to narrowly described security needs.

Anthropic has not yet said whether its detector will satisfy those expectations in a way ordinary affected people can use. Nor has it answered the more important institutional questions.

Will a detector show a probability, a binary label, or an explicit inconclusive state? Will it reveal which passage carries the signal? Will it report the model families and languages it supports? Will it explain that a positive result may reflect proofreading? Will a person accused of misconduct be able to inspect a signed result? Can they reproduce the test? Can they challenge it when the underlying document is unavailable or when a quotation from marked text has been incorporated into their work? Will schools and employers be told, in language impossible to miss, that the result is not proof of authorship or wrongdoing?

These choices determine what the technical system becomes.

A laboratory metric can express uncertainty. Bureaucracies are less patient. A score becomes a flag; a flag becomes a meeting; a meeting becomes a burden placed on the accused to explain a process they cannot inspect. Once a detector is integrated into learning software, publishing systems, freelance marketplaces, applicant screening, or content moderation, the distinction between “signal present” and “violation committed” can vanish between API fields.

Anthropic does not need to intend that outcome to be responsible for anticipating it. Safety engineering is supposed to account for foreseeable misuse. A detector deployed without contestability is not merely incomplete. It externalizes its uncertainty onto the person with the least power in the transaction.

Questions Anthropic Has Not Yet Answered

  • What algorithm or class of algorithm does Claude use to watermark text?
  • Which currently available Claude models are marked today?
  • What are the false-positive, false-negative, and abstention rates by text length, language, model, temperature, and domain?
  • How does the system perform on source code, mathematics, quotations, citations, legal language, and constrained factual responses?
  • Does routine proofreading or translation receive a mark despite the EU editing exceptions?
  • Will users be notified in-product before receiving marked output?
  • Can customers outside the EU opt out?
  • Who will receive detector access, and under what safeguards?
  • Will results identify confidence, location, tested length, model compatibility, and known limitations?
  • Will uploaded text receive zero-retention and privacy-preserving treatment?
  • Can an affected person obtain a signed result, reproduce the test, and appeal an adverse inference?
  • Will Anthropic prohibit use of its detector as the sole evidence for academic, employment, publishing, or professional sanctions?
  • Who will independently audit resistance to removal, mixing, translation, reverse-engineering, and spoofing?
  • How long will detection remain available for retired models and legacy content?

C2PA deserves a separate argument

Anthropic’s file-provenance plan should not be collapsed into the critique of statistical text watermarking.

C2PA Content Credentials are cryptographically signed assertions about a digital asset’s history. A valid credential can help establish that a known signer made particular provenance claims, that the credential is associated with the asset, and that the signed information has not been altered undetectably. That can be genuinely useful for newsrooms, creators, investigators, and distribution platforms.

But C2PA’s own explainer is admirably restrained about what follows. Valid provenance data does not tell a viewer that an image is true. A credential can be valid while the content is misleading. Provenance can be incomplete. An asset without Content Credentials is not automatically untrustworthy. C2PA’s harms guidance expressly warns against that two-tier inference.

Anthropic likewise notes that file metadata can disappear through conversion, resaving, screenshots, and other operations. That is a limitation, not proof that C2PA is pointless. A tamper-evident receipt remains useful even though someone can throw the receipt away. It becomes dangerous only when possession of the receipt is confused with truth, or its absence with fraud.

Text watermarking needs the same semantic discipline.

The steelman is stronger than the sales pitch

There is a compelling case for provenance infrastructure.

Generative systems can produce phishing messages, spam, fake reviews, propaganda, impersonation, and low-cost synthetic material at a scale human investigators cannot match. As model outputs become more fluent, prose-style classifiers lose a stable basis for distinguishing human from machine text. A provider-applied mark can therefore be more principled than asking a classifier whether a passage “sounds like AI.” It may support investigations, measurement, platform enforcement, and deterrence even if it does not catch every adversary.

No security measure has to be perfect to be useful. Door locks can be picked. Email authentication does not eliminate phishing. A watermark that raises the cost of industrial-scale abuse, helps a journalist authenticate one source, or lets researchers identify a large corpus of unedited model output may justify itself.

Nor is Anthropic uniquely imposing an eccentric scheme. Google has already deployed text watermarking in consumer systems. The EU Code was developed through a multi-stakeholder process, and roughly 190 organizations signed one or both sections before the rules took effect. The objective of giving people reliable context about synthetic content is legitimate. NIST’s overview of synthetic-content transparency treats watermarking, provenance tracking, labeling, and detection as complementary tools rather than a single cure.

This is precisely why Anthropic’s thin announcement is inadequate. The company is not introducing a harmless novelty. It is participating in the construction of an attribution layer for digital language. Such a layer deserves more than promises that the mark is invisible, quality-preserving, and detectable later.

The strongest defense of watermarking establishes technical utility. It does not settle interpretation, scope, proportionality, access, or due process. Those are not objections from people who oppose transparency. They are the conditions under which transparency deserves trust.

Publish the receipts before exporting the verdict

A defensible Claude watermark regime would begin by naming the modest claim it can support: a specified detector found evidence of a supported Claude mark in this sample, under these conditions, with this confidence. Every step beyond that sentence requires additional evidence.

Anthropic should publish its evaluation methodology before its detector is used in consequential settings. Results should be broken down by model, language, domain, passage length, sampling configuration, and transformation. Code deserves its own benchmarks; a watermark that leaves ordinary prose unchanged may behave differently when token choice is constrained by syntax, APIs, exact quotations, or mathematical notation. The company should publish both error rates and abstention rates, because a detector that achieves attractive accuracy by declining difficult cases tells a different story from one that classifies everything.

The detector interface should separate “mark detected,” “mark not detected,” and “insufficient evidence.” It should show the tested length, supported model range, confidence, detected region where technically possible, and the class of evidence used. Its most prominent warning should say that the result does not establish authorship, plagiarism, policy violation, or misconduct.

Detection should be locally executable or privacy-preserving wherever feasible. When text must be uploaded, Anthropic should commit to the Code’s zero-retention model, publish security controls, and make signed results available to the person whose work is being judged as well as to the institution judging it.

There must also be an appeals architecture. Anthropic can condition detector access or its terms of use on a rule that no result may serve as the sole basis for an academic, employment, publishing, or professional sanction. It can provide a reproducible evidence package, explain version compatibility, preserve detection for retired models, and create a channel for disputed findings. It cannot prevent every bad inference, but it can refuse to design for frictionless accusation.

Most importantly, Anthropic should respect the difference the law already recognizes between substantial generation and standard assistance. If reliable inference-time scoping is impossible, users outside jurisdictions that mandate marking should receive a meaningful opt-out. All users should be told before marked output is generated. A hidden compliance mechanism should not also be a hidden product feature.

Finally, independent researchers must be able to test the system. That means access sufficient to evaluate removal, mixing, translation, reverse-engineering, collision, and spoofing without exposing keys that would trivially defeat the mark. “Trust us; documentation is coming” is not an acceptable foundation for infrastructure whose errors may be litigated in classrooms and workplaces long before they reach a court.

Society needs better signals about synthetic media. Anthropic is shipping one while leaving downstream institutions to invent its social meaning.

That is how a narrow technical trace acquires powers of accusation it was never capable of earning. Until Anthropic publishes the missing evidence and builds due process around the detector, Claude’s invisible watermark should be treated exactly as the company’s caveats describe it: a clue that Claude may have touched the text, never a verdict on the human who stands behind it.

Methodological note

This article distinguishes three evidentiary categories:

  • Confirmed Anthropic policy: Statements taken from Anthropic’s August 10, 2026 help-center announcement, checked on August 11. The current 31-block article body matched the earliest Internet Archive capture from August 10, and no separate Anthropic detector specification or performance documentation was identified during this review.
  • Class-level technical evidence: Findings from published research on SynthID-Text and other watermark schemes. These establish capabilities, trade-offs, and credible attack classes for text watermarking generally; they are not tests of Claude’s undisclosed implementation.
  • Unresolved Claude-specific questions: Implementation details, performance metrics, access rules, and safeguards that Anthropic had not publicly documented as of August 11, 2026. Their absence is reported as an information gap, not as evidence that Claude’s watermark has a particular defect.

Legal discussion describes Article 50, the official implementation guidance, and the July 2026 Digital Omnibus amendment. It is not legal advice.

Sources used

  1. Anthropic, “How Claude marks AI-generated content”, updated August 10, 2026; accessed August 11, 2026.
  2. Internet Archive, earliest captured version of Anthropic’s announcement, captured August 10, 2026.
  3. European Commission, “Strong backing for the Code of Practice on Transparency of AI-generated Content”, updated August 5, 2026.
  4. EU AI Act Service Desk, Article 50: Transparency obligations for providers and deployers of certain AI systems.
  5. European Commission, Code of Practice on Transparency of AI-generated Content, final PDF, June 2026.
  6. European Commission, Guidelines on the implementation of Article 50 transparency obligations, July 2026.
  7. European Union, Regulation (EU) 2026/1744, Digital Omnibus on AI, July 8, 2026.
  8. Dathathri et al., “Scalable watermarking for identifying large language model outputs”, Nature 634, 2024.
  9. Google DeepMind, “Watermarking AI-generated text and video with SynthID”, May 14, 2024.
  10. Rastogi and Pruthi, “Revisiting the Robustness of Watermarking to Paraphrasing Attacks”, EMNLP 2024.
  11. Sadasivan et al., “Can AI-Generated Text be Reliably Detected? Stress Testing AI Text Detectors Under Various Attacks”, Transactions on Machine Learning Research, 2025.
  12. An et al., “DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation”, EACL 2026.
  13. C2PA, Content Credentials Explainer, version 2.2.
  14. C2PA, Harms Modelling, version 2.4.
  15. NIST, “Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency”, NIST AI 100-4, updated April 8, 2026.