How to Price an AI Chatbot for a Client
Three real pricing structures for an AI chatbot project, flat build fee, build fee plus retainer, and usage passthrough, with the costs each one needs to account for and which client size fits which.
There are three workable ways to price an AI chatbot for a client: a flat one-time build fee, a build fee plus a monthly retainer, or a usage-passthrough model where the client pays their own API bill directly. Which one fits depends less on the chatbot's complexity and more on how predictable the client's usage will be and how much ongoing tuning the bot will need after launch. A simple FAQ bot with light traffic usually wants a flat fee. A bot handling variable, high-volume conversations usually wants usage-passthrough or a retainer that accounts for it.
The three pricing structures, side by side
Before the numbers, it helps to see what each structure actually asks the client to carry, because that's the real difference between them, not the label.
1. Flat one-time build fee
You quote a single number for design, build, and launch. The client pays once, owns the result, and either runs it themselves or comes back separately for changes. As an illustrative example only, not a market figure, a simple FAQ bot answering questions from a fixed knowledge base might run $1,500 to $4,000 as a flat build, while a bot integrated with a booking system or CRM might run $5,000 to $15,000. Your actual number depends on your market, your speed, and what the client's alternatives cost.
What you need to account for inside that number: your build time, the LLM API cost of testing and iterating during development, and a buffer for the inevitable round of revisions after the client actually uses it. You are not accounting for their ongoing API usage at all, that becomes their problem the day you hand over the keys, unless you've separately quoted hosting or a support add-on.
2. Build fee plus monthly retainer
You charge a smaller upfront build fee, then a recurring monthly amount that covers hosting, monitoring, prompt tuning, and a set number of change requests. As an illustrative example, a build fee of $1,000 to $3,000 paired with a retainer of $200 to $800 a month is a plausible range for a small business chatbot, again as a worked example, not a claim about typical market rates.
This structure needs you to account for three ongoing costs inside the retainer: the LLM API usage itself (which you're now absorbing and need to estimate from expected conversation volume), hosting and any vector database or logging infrastructure, and your own time for prompt tuning as the client's product, FAQs, or edge cases shift. Underpricing the retainer relative to actual API usage is the single most common way builders lose money on this model, because usage tends to grow exactly when the bot is working well.
3. Usage-passthrough
The client sets up and pays for their own LLM API account (their own OpenAI, Anthropic, or other provider billing), and you charge separately for the build and, optionally, a smaller monthly fee for maintenance and monitoring that excludes API costs entirely. There's no illustrative dollar range needed here because the API cost simply passes through to the client's own bill, your fee only covers your labor.
This is the cleanest structure for cost accountability, the client sees exactly what their usage costs on their own provider dashboard, with no markup dispute possible. What you still need to account for is your build fee, a maintenance fee if you're offering one, and the practical reality that you'll be doing API key setup and account configuration work for a client who may not be technical enough to do it themselves.
Which structure fits which client
Flat fee fits a client with low, predictable traffic and no appetite for a recurring bill, a local business, a solo professional, anyone who wants to pay once and move on. It also fits you when you don't want the obligation of ongoing support baked into the relationship.
Retainer fits a client who wants the bot to keep improving, expects usage to grow, or simply doesn't want to think about API billing, hosting, or prompt drift themselves. It also fits you better financially if the bot's usage is genuinely hard to predict at quote time, since the retainer gives you room to adjust pricing at renewal instead of eating a bad flat-fee estimate.
Usage-passthrough fits a client who is cost-sensitive about API spend specifically, wants full visibility into what they're paying for, or is already technical enough to manage a provider account. It also fits any client whose usage could spike unpredictably, a viral moment, a seasonal rush, since neither side is stuck absorbing a surprise bill that was based on someone else's guess.
The costs every structure has to account for somewhere
Regardless of which structure you pick, three costs exist and someone has to carry each one explicitly, in the price or in the contract, rather than leaving it to be discovered later.
LLM API usage. Estimate this from expected conversations per month times average tokens per conversation, not from a vague guess. If you can't estimate it, that's itself a signal to lean toward usage-passthrough rather than quoting a flat number you might regret.
Hosting. Wherever the bot's backend, logs, and any vector store live, someone pays for that server or platform every month, and it doesn't disappear just because you charged a flat build fee.
Ongoing maintenance and prompt tuning. Client products change, FAQs get updated, the bot starts giving a wrong answer that needs a prompt fix. This is real, recurring labor even when nothing is technically broken, and it's the cost most new builders forget to price for at all.
A simple way to decide
Ask two questions before you quote. First, how predictable is usage going to be, low and steady points toward flat fee, growing or spiky points toward retainer or passthrough. Second, how much ongoing tuning will this bot realistically need, a static FAQ bot needs almost none, a bot tied to a changing product or inventory needs regular attention, which argues for a retainer that actually pays you for that work rather than a flat fee that quietly turns into free support.
If you're weighing pricing models more broadly, our comparison of per-seat versus usage-based pricing for an AI product covers the same predictability question from a product-pricing angle rather than a client-services one.
Setting expectations before you quote
Whichever structure you choose, put it in writing before work starts: what's included in the fee, what counts as a change request versus a new feature, and who owns the API account. A client who thinks a flat fee includes unlimited future tweaks will be unhappy in month two, and a client on usage-passthrough who doesn't understand their own provider bill will blame you for a cost spike that had nothing to do with your work.
Clear milestones help with this regardless of structure. Our piece on payment milestones for an AI project that actually work walks through how to split a build fee into stages so you're not carrying all the delivery risk upfront, and if the client wants ongoing work after launch, how to price post-launch support for an AI-built app goes deeper into structuring that specific conversation.
For the broader picture of how builders charge for AI work beyond chatbots specifically, see our roundup of AI monetization strategies for builders.
For a version of this same decision broken down by per-conversation billing instead of passthrough, how to price AI chatbot services for a client covers the same three-way tradeoff with a worked pricing table.
That figure covers the initial build. Once the chatbot is live, ongoing upkeep is its own line item, see how to structure a retainer for ongoing AI app maintenance for how to price that part separately.
Frequently asked questions
Should I mark up the client's API costs?
If you're on a retainer where you're absorbing API costs, you're effectively building a margin into the monthly fee already, so a separate markup is unnecessary and can look like double-charging. On usage-passthrough, the client pays their provider directly and you charge only for your labor, so there's no API cost to mark up in the first place.
What happens if API costs rise after I've quoted a flat fee?
On a flat one-time build, this typically isn't your problem since the client owns the result and its ongoing costs after handoff. On a retainer, build in a clause letting you revisit pricing at renewal, or if you set the retainer without a usage cap, expect to eat rising costs until that renewal date.
Is a retainer always better than a flat fee for the builder?
Not necessarily. A retainer creates recurring revenue but also a recurring obligation, you're on the hook for support every month whether or not the client needs much. A flat fee is simpler and lets you move on, which can be the better trade for a small, low-maintenance bot.
How do I estimate API usage costs before quoting?
Estimate expected conversations per month, multiply by average tokens per conversation including any context you're feeding the model, and check that against your chosen provider's current published pricing. Build in a buffer, since early usage estimates from a client are almost always lower than what actually happens once the bot goes live.
Can I switch a client from one pricing structure to another later?
Yes, and it's common. A client that started on a flat fee often moves to a retainer once they realize they want ongoing tuning, and a client on a retainer with unexpectedly high usage may prefer to switch to passthrough so they can see and control their own API spend directly.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


