Pretraining Data Can Be Poisoned through Computational Propaganda
Merged summary
TL;DR - This paper shows that LLM pretraining corpora can be poisoned at web scale by injecting adversarial content through public discussion interfaces, and introduces a method to estimate whether such content survives web crawling and data curation.
- Argues prior poisoning work relied on narrow sources (e.g., Wikipedia) that don't reflect the scale/heterogeneity of real pretraining data and ignored curation pipelines.
- Demonstrates public discussion interfaces as a viable web-scale content injection vector for poisoning attacks.
- Introduces HalfLife, an analysis for estimating adversarial content inclusion in web-crawl-based training data.
- Establishes third-party webpage content as a plausible attack surface for LM pretraining, stressing the need to verify whether poison injections actually reach training data.
Sources (1)
Pretraining Data Can Be Poisoned through Computational Propaganda
TL;DR - This paper shows that LLM pretraining corpora can be poisoned at web scale by injecting adversarial content through public discussion interfaces, and introduces a method to estimate whether such content survives web crawling and data curation.
- Argues prior poisoning work relied on narrow sources (e.g., Wikipedia) that don't reflect the scale/heterogeneity of real pretraining data and ignored curation pipelines.
- Demonstrates public discussion interfaces as a viable web-scale content injection vector for poisoning attacks.
- Introduces HalfLife, an analysis for estimating adversarial content inclusion in web-crawl-based training data.
- Establishes third-party webpage content as a plausible attack surface for LM pretraining, stressing the need to verify whether poison injections actually reach training data.