How to Explain AI Costs to a Client
Setup, per unit, ceiling. Three numbers a client can approve, and the twenty-sample method that gets you to them honestly.
The fastest way to explain AI costs to a client is to stop explaining tokens. Clients do not buy tokens. They buy outcomes, and they need to know three numbers: what it costs to set up, what it costs per unit of work, and what stops the bill from surprising them. Give them those three and the conversation ends. Give them a lesson on context windows and it does not.
Every awkward version of this conversation comes from the same place: the person building understands the cost model and the person paying does not, so the builder explains the mechanism instead of the price. Mechanism is not reassurance.
The three-line model
Put this in the proposal, in the client's own units.
Line | What it covers | Example wording |
|---|---|---|
Setup | Build, integration, prompt and evaluation work, testing | £4,800 one off |
Per unit | The variable cost of the work itself | £0.11 per processed invoice |
Ceiling | The cap that protects them | Hard stop at 6,000 invoices per month, £660 |
Three lines. A client can approve that in a meeting. They cannot approve "roughly $3 to $12 per million input tokens depending on model, plus output tokens at a higher rate, plus caching discounts."
Note what the third line does. It converts an open-ended variable cost into a bounded one, which is the actual thing making the client nervous. Most of the resistance you meet is not about the amount, it is about the absence of a ceiling.
Pick the client's unit, not the vendor's
The per-unit line only works if the unit is something the client already counts. Invoices, tickets, listings, applicants, calls, documents, properties. If they have a number for it in their existing reporting, use that number.
To get from model pricing to their unit, run the job on twenty real examples and measure. That is the whole method:
Take twenty representative inputs from their actual data, including two ugly ones.
Run the finished pipeline and record total input and output tokens per item.
Multiply by current published rates. Vendor pricing is published, for example Anthropic's pricing documentation, and it is quoted per million tokens.
Take the mean, then quote the 90th percentile.
Quoting the mean is how people lose money. The distribution has a tail: the 40-page contract, the invoice with 200 line items, the support thread with 60 messages. Price the tail, and the average case becomes margin rather than a shortfall.
Explaining variability without sounding evasive
Clients hear "it depends on length" as "we do not know what this costs." Reframe it as something they already understand from their own business.
Printing is the analogy that lands most often: a one-page letter and a fifty-page report cost different amounts to print, and nobody finds that mysterious or dishonest. AI processing bills the same way, by how much text goes in and comes out.
What you must not do is present variability as a reason you cannot commit. If you have run the twenty samples, you can commit. Say the number, state the assumption it rests on, and name the ceiling.
Based on your last 200 invoices, this runs at about 11p each. Roughly one in twenty is unusually long and costs about 30p. I have priced at 14p to absorb that. If volume goes past 6,000 a month we should talk, because the economics change in your favour and I would want to renegotiate down.
That last sentence does more for trust than any amount of technical explanation. It signals you are pricing the work, not the confusion.
Three costs people forget to quote
The per-call model cost is usually not the largest number in the first year.
Rework and evaluation. Getting output quality from acceptable to reliable is where the hours go. It is real work, it is invisible in a token calculator, and it belongs in the setup line rather than absorbed silently.
Retries and failed runs. Anything that fails validation and runs again costs twice. On a well-built pipeline this is a few percent. On a rushed one it can be 30 percent, and it will not show up in your sample of twenty unless the ugly examples are genuinely ugly.
The human in the loop. If a person reviews output before it goes to a customer, that person's time is part of the cost of the system. Clients are far happier hearing this up front than discovering later that the automation still needs 20 minutes a day of somebody's attention.
Quoting these three is what separates a proposal that survives month three from one that becomes an argument.
When they ask why it is not cheaper
Two versions of this question exist and they need different answers.
If they mean "the model API looks cheap, why is your invoice not," the answer is that model calls are one input to a working system, alongside integration, error handling, evaluation, and the ongoing job of keeping it correct when the model changes underneath. The cost of the raw call has fallen repeatedly and will keep falling. The cost of making it reliable in their business has not.
If they mean "you used AI, so it took you less time, so charge me less," that is a different conversation entirely, and it is worth having deliberately rather than defensively. There is a full treatment in what to do when a client wants a discount because you used AI.
The conversation when the bill goes up
At some point the number moves against you: volume grew, the workload got heavier, or a model you depend on repriced. How you open that conversation determines whether it is routine or damaging.
Lead with the cause and the fix, not the apology. "Your volume went from 3,100 to 5,400 invoices, which is why the variable line moved" is a good conversation, because it is their success. "Costs have gone up" without a cause invites the suspicion that you mispriced.
Bring one number they can act on: the per-unit cost. If per-unit held steady and only volume grew, the system is working exactly as quoted and there is nothing to defend. If per-unit itself rose, you owe them an explanation of why and what you are doing about it, which is usually caching, a smaller model for the easy cases, or trimming a prompt that grew during the project.
The ceiling is what makes all of this survivable. A client who agreed to a cap is never surprised, only informed, and the conversation is about raising a known limit rather than about a bill that arrived larger than expected. That is the entire reason the third line exists.
Put the numbers where they can see them
Send a monthly line showing units processed and cost, next to the ceiling. Not a dashboard they have to log into, a number in an email. Two consequences follow.
First, no surprises, which is most of what clients want from a vendor. Second, the moment the volume grows, the value of the system is documented in their inbox rather than argued for in a renewal meeting.
FAQ
Should I mark up the model costs?
Either pass them through transparently and bill your work separately, or bundle everything into one per-unit price. Both are honest. What causes trouble is a bundled price that you later describe as "just passing on costs," because it invites a line-by-line audit you did not prepare for.
What if the client asks to see the token numbers?
Show them. Bring the twenty-sample measurement. Curiosity here is a good sign, and the numbers support your quote rather than undermining it, provided you priced the tail rather than the average.
How do I handle prices changing mid-project?
Quote in the client's unit and hold that price for a defined term, usually three to six months. Model prices move in both directions, and absorbing that variance is part of what they are paying you for. Revisit at renewal.
What is a reasonable buffer?
Price at roughly the 90th percentile of your measured per-unit cost. That absorbs the long tail without inflating the quote to the point where it fails a comparison against a competitor who quoted the mean and will miss it.
More on the commercial side: AI monetization strategies, how much to charge for an AI automation project, writing an AI project proposal, and estimating AI project ROI before you start.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


