How to Price an AI Product With Variable Costs
Traditional SaaS has near-zero marginal cost. AI products do not. That single difference breaks flat-rate pricing in a specific, predictable way.
Traditional software has a marginal cost per user of roughly zero. One more subscriber on a note-taking app costs you a rounding error in storage. That single fact is why flat monthly SaaS pricing became the default: heavy users and light users cost the same, so charging them the same is fine.
AI products break that assumption. Every request costs real money, the cost scales with how much people use it, and usage between customers on the same plan can differ by a factor of a hundred. Price an AI product like SaaS and you have built something where your best customers lose you the most money.
Here is how to work out what to charge instead.
Start with cost per action, not cost per user
Before any pricing model, you need one number: what does a single unit of value cost you to deliver?
Use real published rates. Anthropic's pricing documentation lists Claude Haiku 4.5 at 1 dollar per million input tokens and 5 dollars per million output tokens, with Sonnet and Opus tiers costing more. The same documentation works a concrete example: processing 10,000 support tickets, averaging roughly 3,700 tokens per conversation, comes to about 37 dollars on Haiku 4.5.
That is 0.37 cents per ticket. Now you have something to price against.
Do this arithmetic for your own core action before anything else:
tokens in x input rate = input cost
tokens out x output rate = output cost
+ retries, + failed attempts, + any tool calls
= cost per completed actionTwo things people leave out and then discover on the invoice. Retries: if 8% of calls fail validation and get retried, your real cost is 1.08 times the naive figure. And system prompt overhead: a 2,000-token system prompt sent on every request is often larger than the user's actual input, and you pay for it every single time.
The guide to how tokens are counted covers why these counts are less predictable than they look, particularly for non-English text and code.
Cut the cost before you set the price
Your first cost number is usually two to ten times higher than it needs to be. Fix that before pricing around it, because every lever below drops straight to margin.
Right-size the model. Most production work is narrow and repetitive and does not need a frontier model. Routing simple cases to a cheaper tier and escalating only the hard ones is the single biggest lever available, often five times or more. This is the practical argument for smaller purpose-built models: not that they are better, but that they are sufficient for a large share of the work.
Cache the repeated prefix. If every request carries the same long system prompt or document, prompt caching charges cache reads at a fraction of the standard input rate. Anthropic documents cache reads at 0.1 times the base input price, so a hit costs 10% of a fresh read, paying for itself after one or two reuses depending on cache duration.
Batch what is not urgent. Asynchronous processing carries a documented 50% discount on both input and output at Anthropic, and similar mechanisms exist elsewhere. Overnight reports, bulk enrichment, and backfills should never run at interactive rates.
Stacking these on a workload that suits them can move your cost per action by an order of magnitude, which is often the difference between a viable price point and an impossible one.
Four pricing models, and when each breaks
Model | Works when | Breaks when |
|---|---|---|
Flat subscription | Usage is naturally capped by the workflow | A power user runs 200x the median |
Usage-based | Value maps cleanly to a countable unit | Buyers cannot predict their bill |
Credits | Multiple actions with different costs | Credits become a puzzle customers resent |
Outcome-based | The outcome is objectively measurable | Attribution is arguable |
Flat subscription is what customers want and what your competitors probably offer. It works when the workflow itself limits usage: a tool a recruiter uses per candidate is bounded by how many candidates exist. It fails when nothing bounds it. If your product can be run in a loop, someone will.
If you go flat, you need a fair use ceiling stated up front and enforced in code, not in the terms of service. A soft limit with an upgrade prompt at 3x the median usage protects the model without punishing normal customers.
Usage-based aligns price with cost automatically, which solves your problem and creates the customer's. Unpredictable bills are a genuine obstacle in procurement, and buyers routinely choose a more expensive predictable option over a cheaper variable one. Price per unit that the customer already counts and already values: per document, per candidate, per resolved ticket. Never per token, which is your unit and means nothing to them. Usage-based vs flat-rate AI pricing walks through the tradeoff in more detail if you are still weighing the two.
Credits let you charge different amounts for actions with different costs, which is honest. They also introduce a second currency the customer has to reason about. They work when denominated in something intuitive and priced so that the common action is obviously cheap.
Outcome-based is the strongest position when the outcome is countable and clearly attributable. Charging per resolved support ticket rather than per API call puts you on the same side as the customer. It requires that both parties can agree on what counts, which rules it out more often than people expect.
The hybrid most AI products land on
The structure that keeps showing up, because it survives contact with real usage:
A base subscription that covers the median customer, plus metered overage above a generous included allowance.
Set the included allowance so that 80 to 90 percent of customers never exceed it. They get a predictable bill and never think about usage. The heavy minority, who are also usually the ones getting the most value, pay proportionally. Your cost floor is covered by the base fee even for a customer who uses nothing.
Two rules make this work in practice. Show usage against the allowance in the product, so nobody is surprised. And notify before the overage starts, not on the invoice.
Sanity checks before you commit
Gross margin on the heaviest realistic user. Not the median. Model the customer at the 95th percentile of usage and confirm you are still profitable. If you are not, your allowance is too generous or your costs are too high.
Cost as a share of price. A rough target is inference cost under 30% of revenue for a software business. Above 50% and you are reselling compute with extra steps, which is a real business but a different one with different margins and expectations.
Free tier arithmetic. A free tier is a marketing budget with a variable cost attached. Multiply your cost per action by a realistic abuse case, not a polite one, and decide whether you can afford it. Rate limits and required authentication are not optional here.
Model price changes cut both ways. Per-token prices have trended down, which improves margins over time. They can also rise, and a model you depend on can be deprecated with a migration window. Repricing your product every time a vendor reprices theirs is not viable, so leave headroom rather than pricing to today's exact cost. This dependency is the same structural risk covered in the paths to actually making money with AI apps.
Price on value, floor on cost
Everything above establishes your floor, not your price.
What the customer will pay is set by the alternative. If your tool does in four minutes what a person does in three hours, the comparison is that person's hourly cost, not your inference bill. The same reasoning applies to what building with AI costs versus hiring a developer: buyers compare against the option they would otherwise take. Paying for that work yourself raises a similar question: whether to agree on fixed-scope vs hourly pricing for an AI project before signing off on a rate.
Cost tells you the price below which you should not sell. Value tells you the price you can actually charge. AI products are unusual only in that the floor is not zero and moves with usage, which means you have to calculate it deliberately instead of ignoring it.
Start higher than feels comfortable. Raising prices on existing customers is one of the hardest things to do; discounting is easy. When that time comes, how to raise prices on an AI product without churn covers how to do it without losing the customers you already have.
For services rather than a packaged product, the equivalent pricing lever is often outcome-based rather than a flat price: see how to price an AI automation project by outcome.
Frequently asked questions
Should I charge per token?
No. Tokens are your unit of cost, not the customer's unit of value. Charge per thing they already count, such as a document processed or a ticket resolved, and absorb the token math yourself.
How do I stop one customer from destroying my margins?
Per-account rate limits enforced in code, a stated fair use policy, and monitoring that alerts you when an account crosses a multiple of median usage. Build this before launch, because the first time you need it will be a surprise.
Is a free tier viable for an AI product?
Yes, with strict limits, required authentication, and a cost per free user you have calculated deliberately. Treat it as customer acquisition spend with a hard monthly cap rather than an open-ended offer.
What margin should I target?
Software buyers and investors generally expect gross margins well above 70%, which means keeping inference cost meaningfully under 30% of revenue. Lower is workable but changes how the business is valued and how much you can spend acquiring customers.
How often should I revisit pricing?
Recheck unit costs whenever you change models or prompts, since both move token counts, and revisit the pricing structure itself when your usage distribution shifts. Understanding what actually drives token consumption makes these reviews quick rather than exploratory.
Once you have a price, the next practical step is wiring up billing itself. See how to add payments to an AI-built app
Before settling on a number, settle on the pricing model itself. See subscription vs one-time fee for an AI tool.
Once you have settled on numbers, the way you present them matters almost as much. See how to write a pricing page for an AI product for the anatomy of a page that makes usage-based costs feel predictable.
Getting the price right is only half the job. See how to reduce churn on an AI subscription product for the AI-specific reasons subscribers cancel even at a fair price.
Related: structuring a money-back guarantee for an AI product
Related: how to sunset an AI feature without losing customers
If the product in question is a browser extension rather than a full app, the calculus changes some. how to price a browser extension built with AI covers what's different about that case.
Once you're billing for usage rather than charging a flat price, how to bill clients for AI API usage covers the metering side of that.
Pricing is only half the billing picture. Once you have paying customers, how to handle chargebacks on an ai subscription is the operational side worth planning for before it happens.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


