Who Sees Your Data When an AI Router Picks the Model
Routing services forward your prompts to model providers you never contracted with. The eight questions to ask, and why do you train on my data misses.
Who Sees Your Data When an AI Router Picks the Model
When you call a routing or orchestration service, your prompt does not stop at the company you are paying. The coordinator forwards some or all of it to whichever model it selects, run by whichever provider hosts that model. You have one contract and an unknown number of processors.
This is not an accusation. It is the architecture, and several vendors describe it openly. Sakana's Fugu is a pool of coordinated models behind one endpoint by design. The problem is not that routing exists. The problem is that the question most buyers ask, "do you train on my data", does not cover it.
The chain you are actually agreeing to
A direct model API has two parties: you and the provider. A router has at least four.
You send the prompt.
The routing service receives it, and typically stores or logs it long enough to make a routing decision and to debug.
One or more model providers receive whatever portion of it the coordinator forwards, under their own terms, in their own jurisdictions.
Any infrastructure those providers themselves subcontract to, which you will almost never see named.
Each hop has its own retention period, its own training policy, its own security posture, and its own regulator. Your single contract with the router is what ties them together, or does not.
The questions that actually surface the risk
Most vendor security pages answer the wrong questions confidently. These are the ones worth putting in writing, and the answers tell you a great deal about how seriously the vendor has thought about it.
Question | What a weak answer sounds like |
|---|---|
Name every model provider currently in the pool. | "We work with leading providers" |
Can I restrict my traffic to a subset of them? | "Routing is automatic for best results" |
How will I be told when the pool changes? | "We continuously improve our platform" |
What exactly is forwarded, the full prompt or a portion? | "Only what is needed to serve the request" |
Do downstream providers train on forwarded data? | "We do not train on your data" (answers for the router only) |
Which countries do requests get processed in? | Silence, or a single headquarters address |
What do you retain, for how long, and who can read it? | "Industry standard retention" |
Are downstream providers named in your data processing agreement? | No DPA offered |
Notice the pattern in the weak column. Each is a true statement about the router that says nothing about the chain. "We do not train on your data" is the one to watch, because it is usually accurate and usually irrelevant.
What a good answer looks like
A vendor who has done this properly will have a named sub-processor list, a commitment to notify you before it changes, a data processing agreement that flows obligations down the chain, and an option to constrain routing. That combination is standard in mature software procurement and it is reasonable to expect it here.
If you handle personal data in the EU or UK, this is not a preference. A sub-processor list and a written agreement that binds them are what Article 28 of the GDPR requires of a processor, and "our vendor routes to other vendors" is not an exemption. How to tell if an AI vendor is GDPR compliant covers the wider check.
The stability problem underneath the privacy one
Pool composition is the vendor's main cost lever. When a cheaper model becomes available, adding it to the pool improves their margin, and your outputs change without any action on your side or any version number moving.
So the notification question does double duty. You want notice for compliance, and you want it so that a quality regression has a cause you can point at. Without it you are debugging a change you cannot see, which is an expensive way to spend a week. Keep an evaluation set you can re-run on demand, and treat an unexplained shift in results as a routing change until proven otherwise.
When a router is the right call anyway
Plenty of the time. If your prompts contain no personal data, no customer content and nothing commercially sensitive, the chain matters much less and the cost and coverage benefits are real. Internal tooling over public information is a good fit.
The judgement is about content, not architecture. Sort your traffic before you sort your vendors: route the sensitive paths directly to a provider you have a real agreement with, and let everything else take the cheaper path. Splitting by sensitivity is usually easier than negotiating a perfect contract for all of it, and it is the same instinct behind the main risks of building on AI.
FAQ
Is an AI router less safe than calling a model provider directly?
It is harder to verify, which is not the same as less safe. A well-run router with a published sub-processor list can be in better shape than a direct provider with vague terms. The risk is the opacity, and opacity is a choice the vendor makes.
Does a zero data retention promise cover the downstream models?
Only if it says so. Read it carefully, because these promises are usually scoped to the party making them. Ask explicitly whether it flows down. What is zero data retention explains what the term does and does not mean.
Can I find out which model answered a given request?
Some services return it in the response metadata, others consider it proprietary. Ask before you commit, because without it you cannot correlate a bad answer with a routing decision, and your debugging is guesswork.
What should I do if my vendor will not name their providers?
Treat it as a decision, not an oversight. You can accept it for low-sensitivity traffic and route the rest elsewhere. What you should not do is assume a list exists behind the scenes that someone would hand over if asked properly. How to vet an AI vendor has the full process, and what happens to your prompts after you send them to an AI covers the single-provider case.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


