Fixed-Price AI Quotes When Costs Are Variable
How to quote a fixed price when AI costs are variable: the margin math, the three caps that bound your risk, and the contract clause that holds.
Clients want a fixed price. Your AI costs are a usage meter you do not control. Quoting a fixed price when AI costs are variable is survivable, but only if you bound the variable before you sign, rather than hoping the average holds.
The failure mode is specific and common: you quote from a pilot, the pilot ran on polite inputs, production runs on real ones, and by month three the API bill has eaten the margin on a twelve-month engagement.
Start with the unit, not the project
A project price is a guess. A unit cost is a measurement. Before quoting anything, answer one question: what is the smallest repeatable thing this system does, and what does one of them cost?
For a support triage bot, the unit is one ticket. For a document pipeline, one document. For a sales-email drafter, one draft. Then measure, with real samples, not estimates:
# measure from a real run, not from a token estimate
SAMPLES = [
# (input_tokens, output_tokens) from 50 genuine production-like items
(3120, 480), (2890, 510), (9640, 1180), (3050, 470), (18200, 2240),
# ... etc
]
PRICE_IN, PRICE_OUT = 2.00, 10.00 # dollars per million tokens
costs = [(i/1e6)*PRICE_IN + (o/1e6)*PRICE_OUT for i, o in SAMPLES]
costs.sort()
n = len(costs)
print(f"mean ${sum(costs)/n:.4f}")
print(f"median ${costs[n//2]:.4f}")
print(f"p95 ${costs[int(n*0.95)]:.4f}")
print(f"max ${costs[-1]:.4f}")The gap between median and p95 is the number that decides your pricing, and it is usually much wider than people expect. A triage bot whose median ticket costs $0.04 and whose p95 costs $0.31 does not have a $0.04 cost base. It has a distribution, and the long tail is where retries, pasted email threads and attached logs live.
Quote from p95, not from the mean. If that makes you uncompetitive, the answer is to cap the tail, not to quote the mean and hope.
The three caps
Every variable cost in an AI engagement can be bounded in one of three places. Use all three.
Cap | What it limits | How you implement it |
|---|---|---|
Unit cap | Cost of a single item | Truncate input at a documented length, cap max_tokens on output, refuse items above a size threshold |
Volume cap | How many items per period | The contract states an included volume, e.g. 5,000 tickets a month, with a per-unit rate above it |
Model cap | Price per token | You choose the model, not the client. Reserve the right to route to a cheaper one for work that does not need the expensive one |
The unit cap is the one teams skip, and it is the one that saves the engagement. An input truncated at 8,000 tokens cannot cost more than 8,000 tokens of input, whatever the client pastes in. Write the truncation into the spec as a documented behaviour rather than discovering it as an incident.
A worked quote
Concrete numbers, for a six-month support triage build. Measurements from a 200-ticket sample:
Median cost per ticket: $0.041
p95 cost per ticket: $0.310
Client's stated volume: 4,000 tickets a month
Build effort: 9 days
Running costs at p95: 4,000 x $0.310 = $1,240 a month, or $7,440 over six months. At median it would be $984 total, and quoting from that figure would be a $6,456 mistake.
So the quote is built as three lines, not one:
Line | Amount | Basis |
|---|---|---|
Build | $13,500 | 9 days at your day rate, fixed |
Included usage | $1,500 a month | 4,000 tickets at p95 plus a 20% buffer, fixed |
Overage | $0.42 a ticket | p95 unit cost plus 35% margin, variable, billed monthly |
The client gets a fixed number for the thing they are buying and a published rate for the thing neither of you can predict. That is not hedging. It is pricing the two components honestly, and it is far easier to defend in a negotiation than a single number you refuse to break down. Billing clients for AI API usage covers the mechanics of the overage line, and pricing an AI automation project by outcome instead of hours covers the build line when day rates are the wrong frame.
How this differs from the other two pricing questions
Three questions get mixed together in client conversations, and separating them makes all three easier to answer:
Question | What is uncertain | Where it is settled |
|---|---|---|
Fixed scope or hourly? | How long the work takes | |
What is my effort worth now AI does most of it? | What you can defend charging | |
What does it cost to run after delivery? | Volume and input size, which are the client's, not yours | This piece |
Only the third is a pass-through cost. The first two are about your labour, where you control the variable. Conflating them is how people end up absorbing someone else's usage inside a day-rate estimate.
The contract language that matters
Three clauses, and they are short. You do not need a bespoke agreement for this.
A cost ceiling with a notification trigger. State the included volume and that you will notify the client in writing when usage reaches 80 percent of it, before the overage applies. The notification is what keeps the overage from being a surprise invoice, and a surprise invoice is how a profitable engagement becomes a dispute.
A provider price change clause. Model pricing moves in both directions and you do not control it. State that per-unit rates are reviewable with 30 days notice if the underlying provider's published prices change by more than a stated percentage. Vendors do reprice: Google launched Gemini 4 Argon at $2 per million input and $10 per million output as an explicitly introductory rate, per its own announcement.
Model substitution rights. State that you select the model and may change it, provided documented quality thresholds are met. Without this you are contractually locked to today's prices on today's model, which is the worst position available.
On the third clause, define the quality threshold concretely or it is meaningless. A named accuracy measure on a named test set, checked before any substitution, is defensible. The phrase equivalent quality is not.
When a fixed price with variable AI costs is the wrong structure
Some engagements should not have a fixed price, and recognising them early is worth more than any clause:
The client cannot state a volume. If nobody knows whether it is 500 or 50,000 items a month, you are being asked to absorb a hundredfold range.
The input is user-controlled and unbounded, such as a public-facing tool where anyone can paste anything.
The success criteria are still moving. A fixed price on a moving target prices the version you imagined, not the one you will build.
The work depends on a model that does not exist yet. Clients do ask. The answer is a paid pilot, covered in how to run a paid pilot for an AI project.
In all four cases, a short paid discovery that produces the measurement is the right first sale. You cannot price a distribution you have not seen, and the client is usually happier paying for the measurement than receiving a padded number.
If the engagement is already live and underwater
The caps above are preventive. If you have already signed and the usage is above what you priced, you still have moves, in roughly this order of preference:
Find the tail and cap it. Pull the 20 most expensive items of the last month and read them. In most systems the top 5 percent is a recognisable pattern, such as pasted email threads or retried failures, and a single unit cap removes most of the overspend without a conversation.
Cheapen the work that does not need the expensive model. Route the easy majority to a smaller model and keep the expensive one for the hard minority. This is invisible to the client if quality holds, and it is why the model substitution clause matters.
Reopen the volume assumption, with the measurement. Show the client their actual volume against the number they gave you. This is a factual conversation, not a renegotiation, when you bring the data.
Absorb it and reprice at renewal. Sometimes the cheapest option is to eat six weeks of overage rather than spend the relationship capital. Decide this deliberately rather than by drift.
What not to do is quietly degrade quality to protect margin, because the client finds out from their own customers rather than from you.
Questions clients actually ask
Why can you not just give me one number?
Because part of the cost is set by how your customers behave, not by how long the work takes. The build is a fixed number because the effort is yours. The usage is a rate because the volume is theirs. Said that way, most clients find it reasonable.
What if your p95 estimate is wrong?
Then the notification trigger fires at 80 percent and you have the conversation before the invoice, not after. That is the clause's entire purpose.
Can I just add a big buffer and quote one number?
You can, and you will lose competitive deals to people who did the measurement, because your buffer has to cover a tail you have not quantified. A measured p95 plus a published overage rate is almost always a lower headline price than a guessed-at all-inclusive figure.
What happens when model prices fall?
Your margin improves until the client asks for a reduction, which they eventually will. Deciding in advance whether you pass savings on is better than being asked cold, and the same clause that protects you against a rise is what makes the conversation about a fall straightforward.
Does this apply to a retainer rather than a project?
The same three caps apply, with the included volume reset monthly, and the notification trigger matters more rather than less because the client sees twelve invoices instead of one. Pricing an AI agency retainer covers the structure.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


