AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
TL;DR - An empirical forensic evaluation finds that three LLM watermarking methods fail to produce sufficiently robust evidence for court use. Meaning-preserving paraphrasing removed nearly all detected watermarks, while baseline detection reliability was also poor.
- Across 846 valid paraphrase runs, watermark removal reached 100% for KGW and Unigram and 98.3% for SynthID-Text.
- Pre-attack false-negative rates were 70% for KGW, 83% for Unigram, and 80% for SynthID.
- SynthID flagged 5.4% of paraphrased human controls as AI-generated, and 80% of its pristine watermarked outputs fell into an uncertainty deadband.
- None of the tested methods satisfied more than two of five Daubert admissibility factors.