AI watermarking
AI watermarking is a technique that modifies the output of artificial intelligence models, such as large language models (LLMs), allowing it to be later identified as AI-generated. Watermarks are one of several techniques used for artificial intelligence content detection, and include both traditional digital watermarking techniques on media files and metadata, and text watermarking techniques specific to LLMs that have been developed during the 2020s AI boom. Recent research in text watermarking has focused on statistical watermarking techniques, which were first introduced in 2022, with the first production deployments in 2024.
Watermarking has been one focus of the regulation of artificial intelligence, examples including the European Union's Artificial Intelligence Act and the California AI Transparency Act, which both began requiring watermarking of AI output in August 2026.
History
[edit]Early text watermarking research, dating back to 1997, embedded information in existing documents, for example through formatting changes or syntactic and semantic transformations of the text.[1][2] In November 2022, Scott Aaronson first proposed statistical watermarking for LLM outputs.[3] In his scheme, the sampling of successive tokens is steered by a keyed cryptographic pseudorandom function, so that the model's authorship can later be demonstrated statistically by anyone holding the key. A prototype was built at OpenAI by Hendrik Kirchner.[4]
The 2023 paper A Watermark for Large Language Models proposed the "red list" and "soft red list" watermarking methods.[5] This was followed by Google's 2024 paper Scalable watermarking for identifying large language model outputs, which introduced their SynthID method. The paper described the first production deployment of generative text watermarking in Gemini.[6] OpenAI claimed in 2024 that it had built a watermarking system for ChatGPT, but opted not to deploy it out of concerns around false positives for non-native English speakers and the risk of users switching to competitors.[7]
Google announced a SynthID based detector in May 2025,[8] and announced an expanded rollout in May 2026.[9] Anthropic announced in August 2026 that their model Claude would watermark all its output.[10][11]
Regulation
[edit]China announced requirements around AI watermarking in 2022,[12] and the requirement was rolled out in September 2025.[13]
The 2024 European Union Artificial Intelligence Act introduced a requirement for AI model providers to watermark their output, with the provision entering into force in August 2026.[10]
There have been various efforts in the United States to regulate AI watermarking. In July 2023, the Joe Biden administration announced a "voluntary commitment" from seven US AI companies to develop AI safeguards, including via watermarking.[14] Biden's October 2023 Executive Order 14110 identified the development of watermarking systems as a policy goal,[15] but the order was rescinded by Donald Trump in January 2025.[16] Legislative attempts include the federal AI Labeling Act, first introduced in 2023 and reintroduced in 2026.[17] State-level legislation requiring watermarking has been enacted in Washington via HB 1170,[18] and in California via the California AI Transparency Act.[19][20]
Mechanism
[edit]Large language models generate text by producing a probability distribution of the next word, and sampling over that distribution at random. Text watermarks function by modifying the selection process over that distribution. For example, the "red list" method sorts candidate words into a red list, which appears at standard frequency, and a green list, which appears at increased frequency.
Generally, detecting these watermarks require knowledge of a specific secret key known only to the model provider. The presence of watermarks is difficult to detect without knowing the secret key.[21] The ability of a specific piece of watermark detection software to detect watermarks is generally limited to output generated by that provider.[22][8]
Performance and criticism
[edit]Applying watermarks, by definition, changes the outputted text, but views vary on whether it changes the quality of the text. The original Google paper introducing SynthID watermarks concluded that watermarking did not negatively affect user feedback over 20 million responses.[6] However, others observed that watermarking performed worse in edge cases where the output text was short or there were few possible words.[21] Watermarking can cause AI models to fail to adhere to AI guardrails.[23][24]
Rollout of watermarks led to criticism. Some developed tools to remove watermarks from generated text.[10][25] Others considered watermarking to be a step too far, criticising that watermarks would appear in text when LLMs were used for editing and proofreading, contrary to an expectation that watermarks would mark when text was entirely generated by artificial intelligence.[26]
Other approaches
[edit]Statistical watermarking techniques that modify the output probabilities of the models have been the primary focus of research and development, but other approaches to AI watermarking have been proposed or implemented.[22] These include modifying the formatting of the output by substituting similarly displayed Unicode characters or adding non-printing characters,[27] performing lexical substitution on the output by replacing words or phrases with synonyms,[27] or inserting provenance-related metadata into output.[22] Other, non-watermarking, methods of artificial intelligence content detection include using AI classifiers to flag content as AI-generated, or maintaining databases of AI-generated content.[22]
References
[edit]- ↑ Kamaruddin, Nurul Shamimi; Kamsin, Amirrudin; Por, Lip Yee; Rahman, Hameedur (2018). "A Review of Text Watermarking: Theory, Methods, and Applications". IEEE Access. 6: 8011–8028. Bibcode:2018IEEEA...6.8011K. doi:10.1109/ACCESS.2018.2796585. ISSN 2169-3536.
- ↑ Liu, Aiwei; Pan, Leyi; Lu, Yijian; Li, Jingjing; Hu, Xuming; Zhang, Xi; Wen, Lijie; King, Irwin; Xiong, Hui; Yu, Philip (2024-09-03). "A Survey of Text Watermarking in the Era of Large Language Models". ACM Computing Surveys. 57 (2): 1–36. arXiv:2312.07913. doi:10.1145/3691626. ISSN 0360-0300.
- ↑ Aaronson, Scott (2022-11-29). "My AI Safety Lecture for UT Effective Altruism". Shtetl-Optimized. Retrieved 2026-09-03.
- ↑ "OpenAI's attempts to watermark AI text hit limits". TechCrunch. 2022-12-10. Retrieved 2026-09-03.
- ↑ "A Watermark for Large Language Models". 3 July 2023. pp. 17061–17084.
- 1 2 Dathathri, Sumanth; See, Abigail; Ghaisas, Sumedh; Huang, Po-Sen; McAdam, Rob; Welbl, Johannes; Bachani, Vandana; Kaskasoli, Alex; Stanforth, Robert; Matejovicova, Tatiana; Hayes, Jamie; Vyas, Nidhi; Merey, Majd Al; Brown-Cohen, Jonah; Bunel, Rudy; Balle, Borja; Cemgil, Taylan; Ahmed, Zahra; Stacpoole, Kitty; Shumailov, Ilia; Baetu, Ciprian; Gowal, Sven; Hassabis, Demis; Kohli, Pushmeet (2024). "Scalable watermarking for identifying large language model outputs". Nature. 634 (8035): 818–823. Bibcode:2024Natur.634..818D. doi:10.1038/s41586-024-08025-4. PMC 11499265. PMID 39443777.
- ↑ Seetharaman, Deepa; Barnum, Matt (4 August 2024). "Exclusive | There's a Tool to Catch Students Cheating with ChatGPT. OpenAI Hasn't Released It". Wall Street Journal.
- 1 2 Thomson, T.J.; Burgess, Jean; Doyuran, Elif (3 June 2025). Dean, Signe (ed.). "Google's SynthID is the latest tool for catching AI-made content. What is AI 'watermarking' and does it work?". doi:10.64628/AA.edrhjkjdt.
- ↑ "Google, OpenAI announce big expansion of SynthID digital watermarks". Mashable. 19 May 2026.
- 1 2 3 "Claude's new Scarlet Letter watermark is invisible—for now". 13 August 2026.
- ↑ "How Claude's text watermarking works".
- ↑ "Should the United States or the European Union Follow China's Lead and Require Watermarks for Generative AI?". 24 May 2023.
- ↑ "China's social media sites rush to abide by AI-generated content labelling law". September 2025.
- ↑ https://www.pbs.org/newshour/politics/watch-live-biden-announces-ai-safeguards-after-meeting-with-tech-leaders
- ↑ "Key takeaways from the Biden administration executive order on AI". Archived from the original on 1 November 2023.
- ↑ Shepardson, David (21 January 2025). "Trump revokes Biden executive order on addressing AI risks". Reuters.
- ↑ "Senators revive push for AI labels". Politico. 24 June 2026.
- ↑ "Washington will require labels on AI images, rein in chatbots". 3 April 2026.
- ↑ "Top global and US AI regulations to look out for".
- ↑ "AI Regulation in U.S. States: Lessons Learned and Key Takeaways – Communications of the ACM". 19 May 2026.
- 1 2 Smith, Matthew S. (9 September 2026). "How AI Watermarks for Text Balance Clarity and Control". IEEE Spectrum.
- 1 2 3 4 "Detecting AI fingerprints: A guide to watermarking and beyond".
- ↑ "AI model watermarking changes agent behavior".
- ↑ "LLMS respond differently to harmful prompts when AI watermarking is used".
- ↑ Ward, Isabella (19 August 2026). "Coders Say They Already Found Workarounds to Claude's Invisible Watermarks". Wired.
- ↑ "Why Anthropic's AI watermark for Claude text goes further than rivals — for now". Business Insider.
- 1 2 Kumar, Nishant; Singh, Amit Kumar (2025). "Artificial intelligence content detection techniques using watermarking: A survey". Image and Vision Computing. 163 105728. doi:10.1016/j.imavis.2025.105728.