Self-Hosted AI Tool Requests: How to Answer Them
The email usually arrives from your best-paying customer, and it is short: our security team will not approve a hosted tool, can we run this ourselves. There are exactly three honest answers, and t...
Self-Hosted AI Tool Requests: How to Answer Them
The email usually arrives from your best-paying customer, and it is short: our security team will not approve a hosted tool, can we run this ourselves. There are exactly three honest answers, and the difference between them is not technical. It is how much recurring work you are agreeing to do forever, unpaid, in exchange for one contract.
Work out which answer you can afford before you reply, because the reply is very hard to walk back.
The three answers
Say no, and offer the thing they actually need
Most on-premise requests are not requests to run software. They are requests to satisfy a control that somebody wrote down. Ask which specific requirement is failing and you often get an answer you can meet without shipping anything: data residency in a named region, a signed data processing agreement, no training on customer data, a penetration test report, SOC 2 evidence, or a contractual deletion guarantee.
Residency in particular is worth probing, because it is often a proxy for a transfer question rather than a location question. The European Commission's guidance on international transfers of personal data sets out the mechanisms that make a transfer lawful, and a customer whose actual concern is one of those mechanisms can usually be satisfied with paperwork and a region setting rather than with your source code.
Meeting the control is dramatically cheaper than meeting the request. This is the outcome to aim for, and getting there is mostly a matter of having the paperwork ready before the conversation, which is the substance of getting an ai product through a client security review.
Run a single-tenant instance you still operate
You deploy a dedicated instance, in their cloud account or a region they choose, but you keep the keys and the deploy pipeline. They get isolation and residency. You keep control of upgrades, which is the part that matters.
This is the middle option and usually the right one. The cost is real but bounded: your deployment has to become reproducible, your migrations have to run unattended, and you now maintain more than one production environment. Budget for a version skew problem within a year, because a customer will refuse an upgrade window at some point.
License the code and let them run it
They get a build, they run it, you support it. This is the option that feels like the biggest win and is almost always the biggest liability.
The recurring cost is not the deployment, it is everything after: debugging an environment you cannot see, a customer stuck three versions behind, security patches you cannot verify are applied, and your own model provider keys living somewhere outside your control. Reserve this for contracts large enough to fund a person, and be sure that person exists.
What each one actually costs
Option | Up-front work | Recurring cost | Sensible floor |
|---|---|---|---|
Meet the control instead | Documentation, DPA, maybe a region | Near zero | Any paying customer |
Single-tenant instance you operate | Reproducible deploy, unattended migrations | One more environment per customer | A contract worth several times your median |
Licensed self-hosted build | Packaging, install docs, offline licensing | Support for a system you cannot observe | A contract that funds dedicated support |
The floors matter more than the effort estimates. Every one of these decays if you sign it below the level that pays for the ongoing work, and the decay is invisible for about six months. Working out where that floor sits for your business is the same exercise as the rest of ai monetization strategies, applied to a deployment model rather than a price.
The clause that decides it
Before you agree to anything, settle who supplies the model. There are only two answers and they lead to different businesses.
If they bring their own model access, running against their own provider account or a model they host, then a self-hosted deployment is coherent. Costs sit with them, your inference bill does not scale with a customer you cannot meter, and their data genuinely never reaches you.
If they expect your product to keep calling your provider account from inside their network, you have built the worst version of both models: you carry the inference cost, you cannot meter usage reliably, and your API keys sit in an environment you do not control. Say no to that specific shape, every time.
Write the answer into the contract alongside an upgrade obligation, a support boundary that names what you will and will not debug remotely, and a defined end-of-life for versions you no longer patch. Writing an SLA for an ai product covers the response-time half; the version support window is the half people forget.
When the answer is genuinely yes
Some sectors will not move. Defence work, certain health and public sector procurement, and anything covered by a regulator that names data location as a hard requirement. In those cases a self-hosted offer is not a concession, it is the product, and it should be priced as a separate line with its own margin rather than as a discount to get the deal signed.
Two things make it survivable. First, one supported deployment shape, not a per-customer variant, because variants are how a product company turns into a consultancy. Second, an honest position on what runs locally: if the customer also wants the model itself inside their walls, that is a different project with different economics, and whether it is safe to run an ai model locally is the first conversation, not the last.
If what they actually want is your product under their name rather than in their building, you are in a different negotiation entirely, and white labelling an ai tool is the cheaper answer to that question.
FAQ
Should a small team ever offer self-hosting?
Rarely, and never as a discount. If the deal does not fund the ongoing support, the correct answer is the single-tenant instance you operate yourself, which gives the customer isolation without giving you a support surface you cannot see.
What if the customer only wants data residency?
Then deploy to a region rather than to their infrastructure. Residency is usually satisfiable with a hosted deployment in the right jurisdiction plus a data processing agreement, which is weeks of work rather than a permanent obligation.
How do I stop a self-hosted customer falling behind on versions?
Contract for it. Name a supported window, say the current release and the two before it, and make security patches non-optional. Without that clause you will be supporting a two year old build during an incident.
Can I charge more for a self-hosted deployment?
Yes, and you should. It is a different product with a different cost structure. Pricing it as a premium tier is more honest than pricing it as the same product in a different place, and it filters out requests that were never serious.
Self-hosting requests are a good problem to have, but they assume you already have paying customers asking. If you're still pre-revenue, start with finding your first paying customer for an AI side project.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


