What Is Data Poisoning in AI?
Data poisoning means planting content that changes what a model does later. Research suggests the number of documents needed stays roughly constant as models get bigger, which changes who can realistically do it.
Data poisoning in AI means deliberately planting content where a model will absorb it, so that the model behaves differently later. It is not an attack on the running system. It is an attack on the material the system learned from, or on the documents it will retrieve at answer time, and it lands long before anyone notices.
The reason this moved from academic curiosity to practical concern is a counting result. The number of poisoned documents an attacker needs does not appear to scale with the size of the model.
The finding that changed the risk calculation
Research from Anthropic's Alignment Science team with the UK AI Security Institute and the Alan Turing Institute found that poisoning attacks require a near-constant number of poison samples rather than a constant percentage of training data. In the accompanying paper, roughly 250 crafted documents were enough to install a backdoor across models from 600 million to 13 billion parameters. Around 100 documents did not reliably work.
The percentage framing is what makes this striking. For a large model, 250 documents is a vanishingly small fraction of the training corpus. If the requirement had scaled with corpus size, poisoning a frontier model would need millions of documents and would be the exclusive business of well-resourced actors. At a few hundred, it is within reach of one determined person with a blog.
The demonstrated attack was narrow: a trigger string that caused the model to emit gibberish, which is a denial of service rather than a subtle manipulation. Whether more useful backdoors follow the same scaling is not established. Treat the result as a lower bound on feasibility, not as evidence that models in production are compromised.
Three paths, and which one is yours
Poisoning is often discussed as one thing. For anyone building on top of models rather than training them, it is three, with very different relevance.
Path | Who is exposed | Can you do anything |
|---|---|---|
Pretraining corpus | The model provider | No, this is their problem |
Fine-tuning data | You, if you fine-tune | Yes, you chose the data |
Retrieval corpus, the RAG index | You, if you use retrieval | Yes, and this is the live one |
Pretraining. Someone plants documents on the open web, hoping a crawler picks them up. You cannot audit this and you cannot fix it. It belongs to whoever trained the model, and it is a fair question to ask a vendor about.
Fine-tuning. If you fine-tune on scraped data, a public dataset, or user submissions, you have taken on the corpus problem yourself at a much smaller scale. Two or three hundred documents is a plausible fraction of a fine-tuning set, not a vanishing one.
Retrieval. This is where most real exposure sits, and it is the one people overlook because it does not feel like training. If your assistant answers from a document store, and anything can write into that store, then anything can write instructions into your model's context. A support ticket, a shared drive file, a scraped page, a customer-submitted PDF.
Why retrieval poisoning is the practical one
Poisoning a retrieval index needs no training run and no scale. It needs one document that ranks for a query your users ask, containing text that steers the answer.
This is the same mechanism as how prompt injection works, arriving through your own data rather than through user input. The distinction matters for defence: filtering the user's message does nothing, because the hostile text enters after that, when your retriever pulls it in.
Practical controls, in order of value:
Know who can write to the index. If the answer is "anyone who emails support", you have a public write endpoint into your model's context. Segment by trust level and mark low-trust sources.
Show sources with answers. A cited answer lets a user notice that a claim came from a document nobody recognises. Uncited answers hide the whole failure mode.
Keep retrieved content out of the instruction position. Retrieved text is data, not orders. Structure prompts so retrieved passages are clearly delimited and the system instructions say the model must not follow directions found inside them. This is imperfect and still worth doing, as covered in defending your own app against injection.
Re-scan on ingest, not just once. Documents get edited after they are indexed. Check on update, not only on first import.
What is worth worrying about, honestly
If you call a hosted model over an API and do not fine-tune, pretraining poisoning is a supplier-assurance question, not an engineering one. Ask about data provenance during procurement, and move on.
If you run retrieval over content that people outside your team can write, that is a live vulnerability today, and it does not require anybody to attack a model provider. It requires somebody to write a document.
And if you are choosing between worrying about poisoning and worrying about your own data leaking outward, the second is usually the more immediate question. That direction is covered in checking whether a tool trains on your data, and both sit inside the wider map of AI risks.
FAQ
Can data poisoning be removed once it is in a model?
Not easily. A backdoor learned in pretraining cannot be located by inspecting weights in any practical way, and retraining from scratch is enormously expensive. This is why the research emphasis is on preventing poisoned data from entering, and on detecting anomalous behaviour after deployment.
Would I notice if a model I use were poisoned?
Probably not, because backdoors are designed to stay dormant until a trigger appears. Normal evaluation looks normal. That is the point of the technique and the reason it is difficult to reason about.
Is data poisoning the same as prompt injection?
No, though they can produce similar results. Prompt injection puts hostile instructions into the context of a running model. Poisoning puts them into the material the model learned from or retrieves. Injection is immediate and reversible. Poisoning is persistent.
Should this change which model I use?
Rarely. Every model trained on web-scale data faces the same exposure, so it is not a useful differentiator between providers. It is a reasonable thing to ask a vendor about, and a poor basis for choosing one.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


