Dashboard

Fixed-Price AI Quotes When Costs Are Variable

How to quote a fixed price when AI costs are variable: the margin math, the three caps that bound your risk, and the contract clause that holds.

Manuele Estivo
Manuele Estivo
Growth & SEO Lead
1 October 20261 min read

Clients want a fixed price. Your AI costs are a usage meter you do not control. Quoting a fixed price when AI costs are variable is survivable, but only if you bound the variable before you sign, rather than hoping the average holds.

The failure mode is specific and common: you quote from a pilot, the pilot ran on polite inputs, production runs on real ones, and by month three the API bill has eaten the margin on a twelve-month engagement.

Start with the unit, not the project

A project price is a guess. A unit cost is a measurement. Before quoting anything, answer one question: what is the smallest repeatable thing this system does, and what does one of them cost?

For a support triage bot, the unit is one ticket. For a document pipeline, one document. For a sales-email drafter, one draft. Then measure, with real samples, not estimates:

python
# measure from a real run, not from a token estimate
SAMPLES = [
    # (input_tokens, output_tokens) from 50 genuine production-like items
    (3120,  480), (2890,  510), (9640, 1180), (3050,  470), (18200, 2240),
    # ... etc
]
PRICE_IN, PRICE_OUT = 2.00, 10.00   # dollars per million tokens

costs = [(i/1e6)*PRICE_IN + (o/1e6)*PRICE_OUT for i, o in SAMPLES]
costs.sort()
n = len(costs)
print(f"mean   ${sum(costs)/n:.4f}")
print(f"median ${costs[n//2]:.4f}")
print(f"p95    ${costs[int(n*0.95)]:.4f}")
print(f"max    ${costs[-1]:.4f}")

The gap between median and p95 is the number that decides your pricing, and it is usually much wider than people expect. A triage bot whose median ticket costs $0.04 and whose p95 costs $0.31 does not have a $0.04 cost base. It has a distribution, and the long tail is where retries, pasted email threads and attached logs live.

Quote from p95, not from the mean. If that makes you uncompetitive, the answer is to cap the tail, not to quote the mean and hope.

The three caps

Every variable cost in an AI engagement can be bounded in one of three places. Use all three.

Cap

What it limits

How you implement it

Unit cap

Cost of a single item

Truncate input at a documented length, cap max_tokens on output, refuse items above a size threshold

Volume cap

How many items per period

The contract states an included volume, e.g. 5,000 tickets a month, with a per-unit rate above it

Model cap

Price per token

You choose the model, not the client. Reserve the right to route to a cheaper one for work that does not need the expensive one

The unit cap is the one teams skip, and it is the one that saves the engagement. An input truncated at 8,000 tokens cannot cost more than 8,000 tokens of input, whatever the client pastes in. Write the truncation into the spec as a documented behaviour rather than discovering it as an incident.

A worked quote

Concrete numbers, for a six-month support triage build. Measurements from a 200-ticket sample:

  • Median cost per ticket: $0.041

  • p95 cost per ticket: $0.310

  • Client's stated volume: 4,000 tickets a month

  • Build effort: 9 days

Running costs at p95: 4,000 x $0.310 = $1,240 a month, or $7,440 over six months. At median it would be $984 total, and quoting from that figure would be a $6,456 mistake.

So the quote is built as three lines, not one:

Line

Amount

Basis

Build

$13,500

9 days at your day rate, fixed

Included usage

$1,500 a month

4,000 tickets at p95 plus a 20% buffer, fixed

Overage

$0.42 a ticket

p95 unit cost plus 35% margin, variable, billed monthly

The client gets a fixed number for the thing they are buying and a published rate for the thing neither of you can predict. That is not hedging. It is pricing the two components honestly, and it is far easier to defend in a negotiation than a single number you refuse to break down. Billing clients for AI API usage covers the mechanics of the overage line, and pricing an AI automation project by outcome instead of hours covers the build line when day rates are the wrong frame.

How this differs from the other two pricing questions

Three questions get mixed together in client conversations, and separating them makes all three easier to answer:

Question

What is uncertain

Where it is settled

Fixed scope or hourly?

How long the work takes

Fixed scope vs hourly pricing for an AI project

What is my effort worth now AI does most of it?

What you can defend charging

Quoting a project when AI does most of the work

What does it cost to run after delivery?

Volume and input size, which are the client's, not yours

This piece

Only the third is a pass-through cost. The first two are about your labour, where you control the variable. Conflating them is how people end up absorbing someone else's usage inside a day-rate estimate.

The contract language that matters

Three clauses, and they are short. You do not need a bespoke agreement for this.

  1. A cost ceiling with a notification trigger. State the included volume and that you will notify the client in writing when usage reaches 80 percent of it, before the overage applies. The notification is what keeps the overage from being a surprise invoice, and a surprise invoice is how a profitable engagement becomes a dispute.

  2. A provider price change clause. Model pricing moves in both directions and you do not control it. State that per-unit rates are reviewable with 30 days notice if the underlying provider's published prices change by more than a stated percentage. Vendors do reprice: Google launched Gemini 4 Argon at $2 per million input and $10 per million output as an explicitly introductory rate, per its own announcement.

  3. Model substitution rights. State that you select the model and may change it, provided documented quality thresholds are met. Without this you are contractually locked to today's prices on today's model, which is the worst position available.

On the third clause, define the quality threshold concretely or it is meaningless. A named accuracy measure on a named test set, checked before any substitution, is defensible. The phrase equivalent quality is not.

When a fixed price with variable AI costs is the wrong structure

Some engagements should not have a fixed price, and recognising them early is worth more than any clause:

  • The client cannot state a volume. If nobody knows whether it is 500 or 50,000 items a month, you are being asked to absorb a hundredfold range.

  • The input is user-controlled and unbounded, such as a public-facing tool where anyone can paste anything.

  • The success criteria are still moving. A fixed price on a moving target prices the version you imagined, not the one you will build.

  • The work depends on a model that does not exist yet. Clients do ask. The answer is a paid pilot, covered in how to run a paid pilot for an AI project.

In all four cases, a short paid discovery that produces the measurement is the right first sale. You cannot price a distribution you have not seen, and the client is usually happier paying for the measurement than receiving a padded number.

If the engagement is already live and underwater

The caps above are preventive. If you have already signed and the usage is above what you priced, you still have moves, in roughly this order of preference:

  1. Find the tail and cap it. Pull the 20 most expensive items of the last month and read them. In most systems the top 5 percent is a recognisable pattern, such as pasted email threads or retried failures, and a single unit cap removes most of the overspend without a conversation.

  2. Cheapen the work that does not need the expensive model. Route the easy majority to a smaller model and keep the expensive one for the hard minority. This is invisible to the client if quality holds, and it is why the model substitution clause matters.

  3. Reopen the volume assumption, with the measurement. Show the client their actual volume against the number they gave you. This is a factual conversation, not a renegotiation, when you bring the data.

  4. Absorb it and reprice at renewal. Sometimes the cheapest option is to eat six weeks of overage rather than spend the relationship capital. Decide this deliberately rather than by drift.

What not to do is quietly degrade quality to protect margin, because the client finds out from their own customers rather than from you.

Questions clients actually ask

Why can you not just give me one number?

Because part of the cost is set by how your customers behave, not by how long the work takes. The build is a fixed number because the effort is yours. The usage is a rate because the volume is theirs. Said that way, most clients find it reasonable.

What if your p95 estimate is wrong?

Then the notification trigger fires at 80 percent and you have the conversation before the invoice, not after. That is the clause's entire purpose.

Can I just add a big buffer and quote one number?

You can, and you will lose competitive deals to people who did the measurement, because your buffer has to cover a tail you have not quantified. A measured p95 plus a published overage rate is almost always a lower headline price than a guessed-at all-inclusive figure.

What happens when model prices fall?

Your margin improves until the client asks for a reduction, which they eventually will. Deciding in advance whether you pass savings on is better than being asked cold, and the same clause that protects you against a rise is what makes the conversation about a fall straightforward.

Does this apply to a retainer rather than a project?

The same three caps apply, with the included volume reset monthly, and the notification trigger matters more rather than less because the client sees twelve invoices instead of one. Pricing an AI agency retainer covers the structure.

How did this land?

About the author

Manuele Estivo
Manuele Estivo

Growth & SEO Lead

Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.