Pretraining Data Can Be Poisoned through Computational Propaganda
Ranking
Overall
80
Content
95
Popularity
45
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper shows that LLM pretraining corpora can be poisoned at web scale by injecting adversarial content through public discussion interfaces, and introduces a method to estimate whether such content survives web crawling and data curation.
- Argues prior poisoning work relied on narrow sources (e.g., Wikipedia) that don't reflect the scale/heterogeneity of real pretraining data and ignored curation pipelines.
- Demonstrates public discussion interfaces as a viable web-scale content injection vector for poisoning attacks.
- Introduces HalfLife, an analysis for estimating adversarial content inclusion in web-crawl-based training data.
- Establishes third-party webpage content as a plausible attack surface for LM pretraining, stressing the need to verify whether poison injections actually reach training data.
Sources (1)
Pretraining Data Can Be Poisoned through Computational Propaganda
Public signals
Hugging Face upvotes 0
TL;DR - This paper shows that LLM pretraining corpora can be poisoned at web scale by injecting adversarial content through public discussion interfaces, and introduces a method to estimate whether such content survives web crawling and data curation.
- Argues prior poisoning work relied on narrow sources (e.g., Wikipedia) that don't reflect the scale/heterogeneity of real pretraining data and ignored curation pipelines.
- Demonstrates public discussion interfaces as a viable web-scale content injection vector for poisoning attacks.
- Introduces HalfLife, an analysis for estimating adversarial content inclusion in web-crawl-based training data.
- Establishes third-party webpage content as a plausible attack surface for LM pretraining, stressing the need to verify whether poison injections actually reach training data.