Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 16, 2026

Pretraining Data Can Be Poisoned through Computational Propaganda

Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith +2 more

Language models learn from enormous piles of scraped web text, and if someone slips poisoned content into that pile, the resulting bad behavior is hard to catch later. Earlier studies mostly poked at Wikipedia, which is neither as large nor as messy as a real training corpus, and they ignored the filtering pipelines that decide what gets kept. This paper looks at a more realistic route: the public comment and discussion widgets scattered across countless sites, which let an attacker inject text at web scale.

The other useful piece is a method called HalfLife that estimates how much injected content survives crawling and curation to reach the training data. That distinction matters, because injecting poison and having it actually make the cut are two different things. Their analysis points to third-party page content as a real attack surface for pretraining.

This draws on the abstract, which frames feasibility more than a finished attack, so read the paper for the actual evidence and its limits.

From the arXiv (cs.AI) abstract

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited…


More Artificial Intelligence papers