OpenAI's EU text watermark loses most of its signal after light editing
A light editing pass that swaps out one in ten words cut OpenAI's new text-watermark detection rate from about 92% to 66% on 400-token passages, and swapping a quarter of the words dropped it to 17%, according to OpenAI's own published evaluation of the tool it is rolling out this month to satisfy EU AI Act requirements.
What
OpenAI published its approach to the EU AI Act's text-provenance rules on October 5, 2026. Starting that day, API customers anywhere in the world can opt in to watermarking for select models; it stays off by default. Over the coming weeks, OpenAI will add an invisible watermark, built on a technique it calls textGrain, to eligible ChatGPT and Codex output for users in the EU specifically, not globally. TextGrain embeds a statistical signal in the model's word choices rather than in metadata, so the signal survives copy-paste but not necessarily rewriting. The company is also opening applications for approved researchers and expert organizations to access its detector, and it published a technical report on the method.
Detection quality depends heavily on length and subject matter. At a 1% false-positive target, OpenAI's detector caught the watermark in about 80% of 200-token passages versus about 95% of 400-token passages for content like psychology questions; math-heavy text scored substantially lower because there is less flexibility in word choice to encode the signal, per the same report. Benchmark scores for its newest frontier model, Astra, showed no meaningful gap with and without the watermark turned on, including 94.44% versus 93.94% on GPQA Diamond and 49.57 versus 49.76 points on the Artificial Analysis Intelligence Index, per the same report.
Detection holds up only until someone edits the output
For anyone treating this as a dependable compliance signal, the fine print matters more than the headline. The watermark is off by default everywhere except the EU rollout, so most API traffic will carry no signal unless a customer actively opts in. Even inside the EU, the protection degrades fast: a routine synonym pass defeats most of the detection within a few hundred words of editing, which is normal behavior for anyone polishing AI-drafted text before publishing it. OpenAI is not pretending otherwise. Its own writeup states plainly that a missing watermark proves nothing about human authorship, since text can be too short, translated, or edited for detection to work. The detector itself also is not publicly available; only approved researchers and expert organizations can apply for access, so platforms, schools, and newsrooms hoping to self-check suspect text cannot yet do so even if they wanted to.
What to watch next
OpenAI said it plans to open-source textGrain and will revisit each part of the rollout as the technology and EU guidance evolve, but gave no date for wider detector access beyond the current researcher-only program. Worth tracking: whether watermarking becomes a global default rather than an EU-only, opt-in feature, and whether the EU's Code of Practice on AI-generated content pushes OpenAI or rivals like Google (whose SynthID watermark OpenAI says textGrain already matches or beats in its own tests) toward broader public detector access.
Sources
- Our approach to EU text provenance rules: OpenAI, October 5, 2026
- OpenAI will start watermarking ChatGPT's text in the EU: TechCrunch, October 5, 2026
- OpenAI is adding text watermarking in ChatGPT and Codex: The Verge, October 5, 2026
