Claude Just Flagged Your Employee’s Best Work as Fake

By George PapazianAugust 15, 202611 min read
AI GovernanceAI TrendsNewsAI Tools
Claude Just Flagged Your Employee’s Best Work as Fake

AI watermarks are coming. But the bias built into AI detection could cost your business its best people. Here’s what small business owners need to know.

Part 1 of 2 on AI watermarks and small business

When Anthropic announced that Claude would start embedding invisible watermarks in every piece of text it generates, the internet did what the internet does. People got angry. The backlash was loud and immediate, flooding Reddit and X within hours.

Most of that anger centered on privacy. On being “caught.” On the principle that a tool you pay for shouldn’t leave fingerprints on your work.

I read through many threads. What struck me wasn’t the volume of the outrage. It was what nobody was talking about.

Everyone was arguing about whether AI watermarks violate user trust. Almost nobody was asking who actually gets hurt when a crude detection system decides your work is fake. That’s the question that matters more. And the answer has a bias problem that should concern every business owner who hires, contracts, or serves a global market.

What the AI Watermark Does (and What It Doesn’t)

Anthropic’s approach works at the model level. When Claude generates text, the system adjusts token probabilities during generation to embed a statistical pattern that’s invisible to readers but detectable by machine. According to Anthropic’s own help center documentation, published August 11, the watermark travels with the text when you copy and paste it elsewhere, and it may survive light editing.

The EU AI Act is the reason this is happening now. Article 50 of the Act requires AI providers to mark AI-generated content in a machine-readable way. Anthropic signed the Code of Practice on Transparency, and the compliance deadline hit August 2, 2026. Every Claude model released after that date embeds the watermark automatically. Users cannot opt out.

For images and files like SVGs and PNGs, Anthropic uses C2PA, an open provenance standard that attaches signed metadata identifying which model processed the file and when. That’s a different mechanism and a more straightforward one. C2PA creates a hash of the metadata that makes tampering detectable. The text watermark is murkier, less established, and generating far more controversy.

Here is the critical distinction that most coverage has missed. Anthropic itself says the watermark only shows that Claude processed a piece of content. It does not confirm authorship. Fortune reported the same observation: the mark sits at the model level and follows Claude’s output everywhere, regardless of how much or how little the tool contributed. Even asking Claude to proofread a single paragraph or translate a sentence could leave a trace.

That gap between “processed by” and “written by” is where the trouble starts.

Processed by AI and written by AI are two different claims. The watermark only proves the first.
Processed by AI and written by AI are two different claims. The watermark only proves the first.

The AI Watermark Bias Nobody Is Discussing

A research team at Stanford, led by James Zou, published a study in the journal Patterns that tested seven major AI detectors against 91 TOEFL essays. Every essay was written by an educated non-native English speaker. No AI was used in any of them. The writers were human. The ideas were theirs. The words were theirs.

The detectors flagged 61% of those essays as AI-generated.

Sit with that number for a moment. Nearly two out of three essays written entirely by humans were classified as machine-produced. The reason is mechanical: the detectors rely on a metric called perplexity, which measures how predictable the language is. Non-native English speakers tend to use simpler vocabulary, shorter sentence structures, and more conventional phrasing. That reads as low perplexity. Low perplexity reads as “machine-like.”

Zou put it plainly: the detectors score based on the sophistication of the writing, and non-native speakers are naturally going to trail their native-born counterparts on that metric.

By contrast, essays from native English speakers were classified correctly almost every time. Eighteen of the 91 TOEFL essays were flagged unanimously by all seven detectors. Not one. Not a few outliers. Eighteen.

Now layer Claude’s watermark on top of that existing bias. A non-native English speaker uses Claude to polish grammar, tighten phrasing, or fix idioms in a business proposal they wrote from scratch. The watermark embeds. A client or employer runs the document through a detector. The detector was already predisposed to flag that person’s writing style. Now it also finds a watermark confirming AI involvement.

That combination amounts to

Checking access…
Share
George Papazian
About the author
George Papazian
Founder & AI Strategy Consultant, Galyx

30+ years of research strategy on projects for Oracle, Cisco, PayPal, and Walmart — now helping small businesses adopt AI that actually delivers.

More about George →
Related posts

Keep reading