Dashboard

How to Add Usage-Based Billing to an AI-Built App

Usage-based billing is only as good as your event count. Here is the append-only pipeline that keeps that count honest, plus how to avoid bill shock.

Steve Jefferson
Steve Jefferson
Developer Advocate
22 September 20261 min read

Flat subscription pricing is simple to build and simple to explain, which is exactly why so many AI-built apps start there and then hit a wall the moment one customer's usage costs more to serve than their subscription covers. Usage-based billing exists to fix that mismatch, but it is a genuinely harder engineering problem than a fixed monthly charge, because now your billing system has to agree with your application about exactly how much of something happened.

Why This Is Different From Normal Billing

A subscription charge is a single fact: does this customer have an active plan, yes or no. Usage-based billing requires your app to count something (API calls, generated tokens, processed documents, minutes of compute) accurately, in real time or near it, and then reconcile that count against what you actually bill, without double-counting or losing events if a request fails partway through. The count is the hard part. Stripe or a metering provider can turn a correct count into an invoice; nothing external can make your count correct.

Pick a Meter That Matches Value, Not Convenience

The most common mistake is metering the thing that is easiest to count instead of the thing that tracks value to the customer. Raw API call count is trivial to log but often decouples from cost and value fast: one call might process a two-page document, another might process two hundred pages. If your underlying cost driver is tokens processed or compute-seconds, meter that directly, even though it means instrumenting deeper into your pipeline than a simple request counter.

Meter

Good fit when

Bad fit when

API request count

Requests are roughly uniform in cost

Request size or complexity varies a lot

Tokens processed

Your cost is driven by LLM usage directly

You also have fixed per-request overhead worth capturing

Compute time

Jobs vary widely in duration

Most jobs are short and the overhead of timing dominates

Seats or active users

Value scales with people, not volume

Heavy and light users pay the same, which frustrates heavy users

The Event Pipeline That Does Not Lie

Build metering as an append-only event log, not a running counter you increment in place. Every billable action writes an event: customer ID, meter name, quantity, a unique idempotency key, and a timestamp, before the action is considered complete. A running counter that increments on success is tempting because it is simpler, but it cannot survive a retry, a partial failure, or a race condition without either double-billing or under-billing, and you will not notice which one is happening until a customer disputes an invoice.

sql
create table usage_events (
  id uuid primary key default gen_random_uuid(),
  customer_id uuid not null,
  meter_name text not null,
  quantity numeric not null,
  idempotency_key text not null unique,
  occurred_at timestamptz not null default now(),
  billed boolean not null default false
);

The idempotency key is what makes this safe to retry. Generate it from something deterministic about the action itself (a request ID, a job ID) rather than a random value, so that if your billing code runs twice for the same underlying event, whether from a retry, a webhook redelivery, or a crash and restart, the second insert fails on the unique constraint instead of creating a duplicate charge.

Reconciling Before You Bill, Not After

Aggregate `usage_events` into a billing period total and report it to your payment provider (Stripe's metered billing, or a dedicated metering platform) on a schedule, but always keep your own aggregate as the source of truth and treat the provider's number as a report you push, not a count you trust blindly. Before closing out a billing period, run a reconciliation: sum your events, compare against what you last reported, and alert on any drift beyond a small tolerance. Drift usually means a webhook was missed, a job failed silently, or a retry created a duplicate that slipped past the idempotency key somehow. Catching that before the invoice goes out costs you an internal Slack alert; catching it after costs you a support ticket and a credit note.

Showing Customers Their Usage Before the Bill Surprises Them

The single biggest driver of usage-billing churn is not the price, it is the surprise. A customer who wakes up to an invoice three times their expected amount will question whether to keep using the product at all, even if the charge is entirely correct. Build a usage dashboard, even a simple one, that shows running spend against the current period, updated at least daily, and set up a threshold alert (in-app and by email) when a customer crosses 80% of their expected or historical spend. This is not a nice UX extra, it is the difference between usage-based pricing that customers tolerate and usage-based pricing that drives them to cancel.

Handling the Free Tier and Overages Together

Most usage-based products still have a free allowance before metering kicks in. Model this as its own meter with a decrementing balance rather than as a separate code path from paid usage, so the same event pipeline and reconciliation logic covers both. When a customer crosses from free allowance into paid usage mid-period, that transition should be visible in their usage history, not silently absorbed into a single number that leaves them unable to tell how much of their bill was allowance overage versus expected use.

Refunds and Disputed Usage

Eventually a customer will dispute a charge, claiming they did not generate the usage you billed for, often after a bug on their side (a retry loop, a misconfigured integration) drove real but unintended volume. This is exactly why the append-only event log matters beyond fraud prevention: when a dispute comes in, you can pull the actual event history for that customer and period, timestamped and itemized, rather than defending a single aggregate number you cannot break back down. Most disputes resolve quickly once you can show, event by event, what happened and when, and a fair number turn out to be the customer's own automation running away, visible clearly in a timestamp cluster no human usage pattern would produce.

Decide your refund policy before the first dispute arrives, not during it. A reasonable default: usage billed is usage that genuinely occurred and is not refundable by default, but a documented bug on your side (metering error, duplicate counting that slipped past reconciliation) is refunded in full, credited automatically once identified rather than requiring the customer to notice and complain first.

Frequently Asked Questions

Should I build metering myself or use a third-party platform?

For a single, simple meter, Stripe's native metered billing is usually enough and avoids building a reconciliation system yourself. Once you have multiple meters, tiered rates, or need usage data inside your own product (like the dashboard above), owning the event log becomes worth the extra work, with the third-party platform as the invoicing layer on top of your own source of truth.

How often should usage be reported to the billing provider?

Daily is a reasonable default for most B2B products, close enough to real time that customers can course-correct before period end, without the operational overhead of streaming every event live. High-usage or enterprise-tier customers may justify more frequent reporting so large overages surface faster.

What happens if my usage event pipeline goes down for an hour?

This is exactly why the append-only event log with idempotency keys matters: your application should keep queuing events (or generating them from an authoritative source like request logs) even if the pipeline that reports to the billing provider is temporarily down, then catch up once it recovers, rather than losing that hour of usage entirely.

Usage-based pricing decisions connect directly to unit economics; see how much does an AI feature cost per user for the cost side of this equation before you set meter rates. If you are building this for a multi-tenant product, how to add multi-tenant support to an AI-built app covers the account isolation this billing system needs to key off correctly, and how to add SCIM provisioning to an AI-built app is the next enterprise-readiness piece most usage-billed B2B apps need. For broader architecture guidance, start at how to build an app with AI.

For the upstream decision of what to meter and how to price it before you build the pipeline below, see how to add usage-based pricing to an AI app.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.