Dashboard

Data Processing Agreement for an AI Product

A data processing agreement for an AI product is the contract that says what you do with your customer's data, which third parties touch it, and what happens when something goes wrong.

Manuele Estivo
Manuele Estivo
Growth & SEO Lead
2 September 20261 min read

A data processing agreement for an AI product is the contract that says what you do with your customer's data, which third parties touch it, and what happens when something goes wrong. If you sell software that sends customer data to an AI provider, you almost certainly need one, and the specific thing that makes yours harder than a normal DPA is that you are not the last stop. Your customer's data passes through you to a model provider, and your DPA has to account for that chain honestly.

This is not the same document as a privacy policy. A privacy policy is a public notice to end users. A DPA is a contract between you and your business customer, and it is what their legal team will actually read.

The chain you have to disclose

The structure is simple once you see it. Your business customer is the controller: it decides why the data is processed. You are the processor: you handle it on their instructions. The AI provider you call is a sub-processor: it processes on your instructions, which came from theirs.

That third link is the part teams get wrong. Founders write a DPA that describes their own handling carefully and never names the model provider, either because they did not think of it or because they would rather not advertise which one they use. Both are a problem. A controller cannot meet its own obligations without knowing who its data reaches, and a DPA that hides the sub-processor is one that fails the first serious review it meets.

The clauses that carry the weight

Most DPA templates are boilerplate that you can adapt. Six clauses are where AI-specific substance goes, and they are the ones worth writing yourself.

Clause

What it must say for an AI product

Sub-processors

Name the model provider, the hosting region, and how you notify customers of changes

Purpose limitation

That customer data is used to serve their requests only, not to improve your product generally

Training

Explicitly, whether any customer data reaches a model training process, yours or the provider's

Retention

Your own retention window and the provider's, stated separately, because they differ

Sub-processor terms

That your agreement with the provider imposes protections no weaker than this one

Deletion

What deletion means when data may sit in a provider's buffer you do not control

The training clause is the one customers care about most, and the one most likely to be read closely. Say it plainly. "Customer data is not used to train any machine learning model, by us or by our sub-processors" is a clear commitment if it is true. If it is only true for some tiers or some features, say which, because an overstated claim here is the kind that gets discovered.

Retention is where the honest answer is awkward

You control your database. You do not control your model provider's storage. Most providers retain API content for a limited window for abuse detection, commonly around 30 days, and that window is invisible to your deletion endpoint.

There are three defensible positions, and picking one is better than being vague:

  1. Disclose the gap. State your retention period, state the provider's, and note that deletion in your system does not immediately purge the provider's buffer. This is accurate and most customers accept it.

  2. Eliminate it. Get a zero data retention agreement with the provider so there is no buffer to disclose. Our explainer on zero data retention covers what that involves and why it is usually gated behind an enterprise tier.

  3. Avoid it. Do not send identifiable data to the provider at all. Redact at your application layer so the question does not arise.

The third is underrated. A great deal of what teams send to a model does not need a name or an account number attached.

You will need a lawyer to sign off eventually. You can get most of the way there first, which makes that review cheaper and shorter.

  • Start from a real template. The European Commission's standard contractual clauses for controllers and processors, adopted in June 2021 to satisfy Article 28(3) and (4) of the GDPR, are a widely accepted starting point.

  • Fill in the specifics before the lawyer sees it. The parts that need a lawyer are the liability and indemnity clauses. The parts that need you are the data flows, sub-processors, retention windows and security measures, and only you know those.

  • Read your provider's own DPA first. Your commitments to your customer cannot exceed what your provider commits to you. Work backwards from their terms rather than promising something and discovering later that you cannot deliver it.

  • Keep the sub-processor list on a public page and reference it from the contract. Then adding a provider is a page update with notice, not a contract amendment with every customer.

Once it exists, the DPA does real commercial work: it is one of the first documents requested in enterprise procurement. Having it ready shortens the process, in the same way described in getting an AI product through a client security review. It pairs naturally with an SLA, and both belong in the standard packet you send buyers, alongside the commercial terms covered in our monetization guide.

FAQ

Do I need a DPA if I only use an AI API and store nothing myself?

If you process personal data on behalf of a business customer, yes, even briefly. Passing data through rather than storing it does not remove the processor relationship.

Is a privacy policy enough?

No. They serve different purposes. The privacy policy is a public notice to individuals; the DPA is a contract with your business customer, and it is what their procurement and legal teams will require.

Do I have to name my AI provider?

In practice yes. Sub-processor disclosure is a standard requirement, and a controller cannot meet its own obligations without knowing where its data goes. Keeping the list on a versioned public page is the least painful way to handle it.

What if my provider changes its retention policy?

Your DPA should reference the sub-processor list and provide for notice of material changes rather than hard-coding a specific window into every customer contract. Check whether the provider is GDPR compliant before you rely on their terms.

How did this land?

About the author

Manuele Estivo
Manuele Estivo

Growth & SEO Lead

Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.