How to Vet an AI Vendor Before You Hand Over Data

Seven questions, the exact wording to use, and what a real answer sounds like next to a marketing one. Including the question most checklists leave out: whose model is actually behind this.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
6 August 20261 min read

Vetting an AI vendor comes down to seven questions, and the useful part is not the questions, it is knowing what a real answer sounds like. "We take security seriously" and "we retain prompts for 30 days, then delete, and here is the clause" are both answers to the same question. Only one of them tells you anything.

Send these in an email rather than asking on a call. You want the answers in writing, and you want to see which ones take three weeks to come back.

1. Do you train on our data, by default?

Ask exactly that, with "by default" included. Plenty of vendors will not train on your data if you fill in a form, which is a different fact from not training on it.

A real answer names the default state and where it is written: "No. Business tier accounts are excluded from training by default, clause 4.2 of the DPA." A weak answer is "your data is never used without your consent," where consent turns out to be the terms you accepted at signup.

Follow up on whether that covers everything or just model training. Some vendors do not train but do retain content for abuse monitoring, human review, or product analytics, which may or may not matter to you but should not be a surprise.

2. How long do you retain prompts and outputs, and can we turn it off?

You want a number of days and a mechanism. "We retain API inputs and outputs for 30 days for abuse monitoring, and zero-retention is available on request for eligible accounts" is a real answer. "We retain data only as long as necessary" is a sentence that survives any subsequent behaviour.

Retention is usually where the practical exposure lives. A model that does not train on your data but keeps it in logs for a year still means your customer records exist on their infrastructure for a year.

3. Who is the actual model provider?

This is the question most vendor checklists skip, and it is the one that changes the answer to all the others.

A large share of AI products are interfaces over somebody else's model. That is a perfectly reasonable way to build a product, and it means your data goes to at least two companies. If the vendor uses a frontier model through an API, their data policy and the model provider's data policy both apply, and the weaker one is the one that governs you.

Ask directly: which models do you use, from which providers, and do you send our content to them. A vendor that cannot answer this quickly is either evasive or does not know, and neither is reassuring. Understanding what a wrapper app is helps you read the answer.

4. Who else touches the data?

The formal version is a subprocessor list, and any vendor selling to businesses should have one on request or on their site. It should name each subprocessor, what they do, and where they operate.

What you are looking for is anything you did not expect: an analytics provider receiving prompt content, a support tool with production access, an offshore annotation team doing quality review. None of these are automatically disqualifying. All of them are things you would rather learn now.

5. Can we delete our data, and how long does it take?

Ask for the mechanism and the elapsed time, including backups. "Deletion via the dashboard, removed from live systems immediately and purged from backups within 35 days" is a real answer that acknowledges backups exist. Anything that claims instant complete deletion is either wrong or is not counting backups.

If you have customers in a jurisdiction with deletion rights, this answer becomes part of your own compliance position, not just your risk appetite. It also needs to match what your own privacy policy for an AI app promises, because you cannot offer your users a deletion guarantee your vendor will not honour.

6. What is your incident history and how do you disclose?

Ask two things: have you had a security incident affecting customer data in the last 24 months, and what is your notification commitment when one happens.

The answer to the first is not disqualifying. Most companies of any size have had something, and a candid account of an incident with what changed afterwards is a better signal than a clean claim. The second answer should be a number of hours in the contract, not a promise to act promptly.

7. What happens when you are acquired or shut down?

Small AI vendors get acquired constantly and shut down almost as often. You want to know whether you can export your data in a usable format, whether the data policy survives a change of ownership, and what notice you get.

The realistic answer for a startup is limited, which is fine as long as you size your dependency accordingly. The bad outcome is discovering at acquisition that your two years of data are exportable only as a screenshot.

Reading the answers

A few patterns worth weighting:

  • Specificity beats reassurance. Dates, clause numbers, day counts and named subprocessors are all checkable. Adjectives are not.

  • A pointer to the contract beats a claim in an email. If it matters, it should be in the DPA, and a good vendor will tell you where. Once you have picked a vendor, how to negotiate a contract with an AI vendor covers how to turn those same answers into binding terms.

  • Speed of reply is data. A vendor that answers these in two days has answered them before. Three weeks and a rewritten marketing page means you are the first person to ask, which tells you something about who else is buying.

  • "It depends on your tier" is fair. Just get the answer for the tier you will actually be on, not the enterprise one.

Sizing the answer to the decision

Not every purchase needs this. Run the full list when the vendor will touch customer personal data, anything under a confidentiality obligation, or anything whose failure would be a material problem for your business. For a tool that summarises public documents, questions one and two are enough.

The general principle is that your exposure is proportional to what you send, not to what the vendor promises, and the cheapest control is still sending less. What to think about before giving AI access to your data covers that side, and the vendor questions here are what you do once you have decided some data has to go.

It is also worth being clear with yourself about where liability lands. A vendor's error that reaches your customer is, from your customer's point of view, your error, and contracts rarely change that impression. Who is responsible when AI makes a mistake goes into that in more detail, and it is the reason vetting is worth doing even when the vendor is much larger than you are.

Vendor exposure is one input into a larger picture. The full map of AI risks covers the categories this checklist does not reach, including the ones that come from your own use of a model rather than from anything a supplier does.

Inflated AI marketing claims are exactly what vendor vetting is meant to catch. See what AI washing looks like for real enforcement cases and a checklist for spotting it.

Vetting only covers tools you know about. See shadow AI for the tools your team adopts without telling you, and what to do about it.

FAQ

Do I need a DPA for every AI tool?

If personal data goes into it, yes, and in several jurisdictions that is a legal requirement rather than a preference. If no personal data is involved, a DPA matters less and the retention and training answers matter just as much. If you are selling into the EU, this is exactly what a GDPR compliance check for an AI vendor needs to confirm.

Is a SOC 2 report enough on its own?

No. SOC 2 tells you a company has controls and that an auditor tested them. It does not tell you whether they train on your data, how long they keep it, or who their subprocessors are. It is a useful baseline and a poor substitute for the seven questions.

What if the vendor will not answer?

Treat non-response as the answer. There is no version of this where a vendor cannot say whether they train on your data and it turns out fine. For a low-stakes tool, proceed with the assumption that everything you send is retained indefinitely and behave accordingly.

How often should I re-check?

Annually for anything material, and immediately on two triggers: a change of ownership, and a change to the model provider behind the product. The second one changes your data path without changing anything you can see from the outside.

Does self-hosting an open-weight model remove the problem?

It removes the vendor from the data path, which is a genuine advantage, and replaces it with responsibility for the infrastructure. That is a real trade rather than a free win, and it sits alongside the other differences between open-weight and closed models.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.