跳到正文
原文
The Decoder· Matthias Bastian·· 4 小时前AI 评分61

OpenAI将为欧盟ChatGPT与Codex文本加textGrain水印,全球API可选

OpenAI will watermark ChatGPT text in the EU but makes it optional for API users worldwide

AI 导读

OpenAI将在未来数周为欧盟的ChatGPT和Codex用户开启textGrain不可见水印,在用词中嵌入统计信号以符合欧盟规则,同时把全球API水印设为可选,不同于Anthropic对Claude的强制方案。

正文 · 原文

OpenAI is adding invisible watermarks to ChatGPT text to comply with EU rules. Unlike Anthropic, it will let API customers worldwide turn them off.

The EU AI Act reportedly requires AI providers to label generated text in a machine-readable format. OpenAI's answer is textGrain, a technology that embeds an invisible statistical signal in the model's word choices. It works much like Claude's SynthID watermark, which uses Google's open-source technology.

OpenAI says it will turn on watermarking for ChatGPT and Codex users in the European Union over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory watermarking for Claude, which applies globally regardless of how users access the models. OpenAI's API watermarking will also be available through cloud partners such as Microsoft Azure in the coming weeks.

Text edits sharply reduce detection rates

OpenAI says textGrain matched or beat other approaches in internal tests, including Google's SynthID for text. The company also plans to release the technology as open source so others can build on it.

Detection rates depend heavily on text length and subject matter. With the detector set to a target false-positive rate of 1 percent, it identified watermarks in about 95 percent of 400-token passages about psychology. At 200 tokens, that fell to about 80 percent.

Detection rates were "substantially lower" for math content, where the model has less freedom to choose words, OpenAI says. Longer passages might strengthen the statistical signal and improve detection, though results could still depend on the content. OpenAI provides no data to test that assumption.

Detection depends heavily on text length and subject matter. At 400 tokens, roughly 300 words, the detector identifies about 94 percent of watermarks in psychology passages but only about 60 percent in math passages. | Image: OpenAI

Editing the text also makes the watermark harder to detect. OpenAI says replacing just 10 percent of words with synonyms cuts detection rates for 400-token passages from about 92 percent to 66 percent. Replacing a quarter of the words drops detection to 17 percent, making the watermark easy to evade. That weakness could work in OpenAI's favor. Users who don't want their ChatGPT use detected might switch to open-weight models if its watermarks become harder to remove.

Even limited editing sharply reduces detection rates. Replacing 25 percent of the words pushes detection below 20 percent, even in 400-token passages. | Image: OpenAI

A technical report explains how textGrain works in depth. Anthropic hasn't published detection rates for Claude's watermark, though it says its system holds up well against edits.

OpenAI says watermarking doesn't hurt output quality, citing tests of its frontier model Astra. The company reports no significant performance differences with watermarking on or off across eight benchmarks, including GPQA Diamond, BrowseComp, and DeepSWE. Those results don't establish whether watermarking affects writing quality, something that critics have also questioned about Claude's watermark.

OpenAI also says a detected watermark reveals nothing about how much human creativity or editing went into the text. It doesn't establish ownership, assign responsibility, identify a user, or show whether the text is accurate and failing to detect a watermark doesn't prove a human wrote the text, either. The passage could be too short, have been edited or translated, or come from a model the detector doesn't support.

OpenAI will limit detector access for now

Only selected researchers and specialist organizations will initially get access to textGrain's detector. They can apply through an application form. OpenAI will grant access case by case under the EU's Code of Practice. Anthropic takes a similar approach with its detection API.

The tool will report only whether it detected an OpenAI watermark. It won't identify users or reveal their prompts or conversations.

OpenAI says it's restricting access because the detector can flag unmarked text or miss watermarks. The company plans to expand access "when we believe results can be interpreted responsibly." It hasn't said when that might happen.

Existing verification tools for images and audio remain publicly available, including openai.com/verify and the Content Provenance API.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

来源:The Decoder · the-decoder.com