The compliance play that reveals the tech's limits
OpenAI just announced its approach to text watermarking under the EU AI Act, and the most interesting part isn't what they're shipping—it's what they're not making public. Starting today, EU ChatGPT users get invisible watermarks on their outputs via a new technique called textGrain. But the detector? That's researchers-only, by application.
This isn't coyness. It's an acknowledgment that text watermarking remains fundamentally brittle technology, and OpenAI is threading a regulatory needle while being unusually transparent about the failure modes.
How textGrain works (and where it breaks)
The textGrain approach embeds "an invisible statistical signal" into the model's word choices during generation. The detector then hunts for that signal to assess whether OpenAI's models touched a given passage. OpenAI claims it matched or exceeded other methods they tested, including Google DeepMind's SynthID for text, in their internal evaluations.
But here's where it gets interesting. Even under ideal conditions, the numbers are sobering:
- At a 1% false positive rate, the detector catches watermarks in ~80% of 200-token passages
- That jumps to ~95% for 400-token passages—if you're working with flexible content like psychology explanations
- For constrained domains like mathematics, detection rates drop "substantially" (OpenAI doesn't specify how far)
The edit resistance is even worse. In 400-token passages, replacing just 10% of words with synonyms tanks detection from 92% to 66%. At 25% replacement, you're down to 17%.
These aren't edge cases. These are Tuesday on the internet.
The phased rollout tells the real story
OpenAI's deployment strategy reveals their own confidence level:
- EU ChatGPT and Codex: Watermarking goes live automatically over the coming weeks (compliance requirement)
- API globally: Opt-in only, off by default (letting customers decide)
- Detector access: Application-required, researchers and "expert organizations" only
That third point is the tell. If the detector were reliable enough for general use, they'd ship a public verification endpoint like they did for images and audio. Instead, they're explicitly worried about "the risk of missed watermarks and false positives" in the wild.
The detector won't identify users, prompts, or conversations. It just answers: did an OpenAI model touch this text? And even that narrow question comes with a long list of caveats.
What a watermark doesn't tell you
OpenAI includes a surprisingly detailed section on what you can't conclude from detection, and it's worth internalizing:
- No measure of human contribution: The watermark can't distinguish between a fully AI-generated essay and a human-edited paragraph that ran through ChatGPT for grammar cleanup
- No ownership or responsibility signal: Detection doesn't establish who wrote it, who owns it, or whether its use was lawful
- No user identification: You can't trace text back to a person or account
- No accuracy verification: The watermark is silent on whether content is true, misleading, or harmful
- No proof of human authorship from absence: Missing watermarks prove nothing—text could be too short, edited, translated, from an older model, or generated elsewhere
That last point deserves emphasis. The edit-sensitivity numbers mean that absence of a watermark is essentially uninformative. Anyone motivated to avoid detection has trivial countermeasures.
The benchmark impact (mostly fine)
To OpenAI's credit, watermarking doesn't appear to tank model quality. Their latest frontier model, Astra, shows minimal performance difference with watermarking enabled across eight benchmarks. The differences are in the noise: 49.57 vs 49.76 points on the Artificial Analysis Intelligence Index, 72.80% vs 71.68% on DeepSWE v1.1.
This matters because it means the technique isn't fundamentally incompatible with high-quality generation. The problem isn't that watermarking ruins outputs—it's that detecting watermarks in edited outputs is nearly impossible.
The regulatory paradox
Here's the tension: The EU AI Act requires making generated text "identifiable in a machine-readable way." OpenAI is shipping technology that technically complies but which they openly acknowledge can't reliably identify AI text once it encounters the real world.
It's compliance theater with a side of radical transparency. They're documenting the technology's limitations more thoroughly than the regulation likely anticipated, while simultaneously rolling it out in a way that minimizes harm from false positives and missed detections.
The researchers-only detector access is actually the responsible call here. Public access would inevitably lead to misuse: teachers accusing students based on unreliable signals, platforms auto-flagging content with high false-positive rates, or governments building enforcement on top of fundamentally uncertain detection.
What happens next
OpenAI plans to open-source textGrain and update their technical report with additional details in the coming weeks. This is good—if text watermarking is going to improve, it needs the research community stress-testing it in adversarial conditions.
They're also working with cloud partners to enable watermarking for OpenAI models accessed through third-party services, which is necessary for any kind of ecosystem-wide provenance story.
But the broader question remains: Is text watermarking even the right tool? Images and audio have different properties—they're harder to edit semantically while preserving meaning, and their provenance matters more for specific harm vectors like deepfakes. Text is infinitely malleable, routinely edited, and often collaborative.
OpenAI describes this as "a phased approach" that reflects "the technology's limitations." That's diplomatic language for: we're not confident this actually works at scale, but the regulation requires us to ship something.
The bottom line
If you're in the EU, your ChatGPT outputs will soon carry invisible watermarks. If you're an API customer anywhere, you can opt in. And if you're a researcher studying provenance, you can apply for detector access.
But if you're hoping watermarks will reliably distinguish AI from human text in the wild? OpenAI's own numbers say: not yet, and maybe not ever. The 66% detection rate after light editing isn't a bug to be fixed—it's a fundamental property of trying to watermark something as fluid and rewritable as natural language.
The most honest thing in this announcement is what's not being shipped publicly. Sometimes the absence speaks louder than the presence.