Dashboard

How to Handle Failed Payments in an AI-Built App

The charge failing is the easy part. What breaks in AI-built apps is the state machine afterwards: access policy, retries, and a definite end to the cycle.

Steve Jefferson
Steve Jefferson
Developer Advocate
20 September 20261 min read

How you handle failed payments in an AI-built app comes down to three things your payment provider will not do for you: deciding what a customer can still access while their payment is failing, retrying on a schedule that respects why it failed, and ending the cycle at a definite point instead of leaving accounts in limbo forever. The charge failing is the easy part. The state machine afterwards is the work.

Most generated payment code stops at the checkout. It handles the happy path, stores a subscription ID, and has no concept of what happens in week three when a card expires.

Why AI-built apps handle failed payments badly

When you ask a model to "add Stripe subscriptions," you reliably get checkout, a webhook handler, and a subscription record. You reliably do not get dunning, because dunning is a business policy rather than an API call, and the model has no way to guess your policy.

The result is a recognisable bug: the subscription goes past_due in the provider, your database still says active, and a customer who stopped paying two months ago still has full access. Nobody notices until you reconcile revenue.

If you have not built the payment layer yet, adding payments to an AI-built app covers the foundation this post assumes, and the full guide to building an app with AI covers where it sits in the wider build.

Separate the two questions the provider conflates

A failed payment has two independent states, and collapsing them into one boolean is the root of most bugs here.

Question

Owned by

Values

|---|---|---|

Is the money collected?

Your payment provider

paid, failed, retrying, abandoned

What can the customer do?

You

full access, restricted, read-only, locked

Providers give you the first. They do not decide the second, and they should not. A B2B tool with annual contracts and a consumer app with 5 a month subscriptions want completely different answers.

Model them as separate columns. billing_status mirrors the provider. access_level is yours, derived from billing status plus grace policy plus manual overrides. When support needs to extend someone's access for a week, they change the second one without lying about the first.

Retry on the reason, not on a fixed timer

Card networks return a decline reason, and the reasons fall into groups that want different treatment.

  • Insufficient funds. Genuinely worth retrying, and timing matters. Retrying around a typical payday pattern rather than in fixed 24-hour increments materially improves recovery.

  • Expired card. Retrying the same card is pointless. This needs a customer action, so the email matters more than the retry.

  • Do not honour. The vaguest and most common code. A small number of retries, then treat it as needing customer contact.

  • Lost or stolen card, or a hard decline. Stop immediately. Retrying a card reported stolen is a fraud signal against your own account.

Your provider's automatic retry logic handles some of this. The part it cannot handle is your messaging and your access policy, which is why you still need the state machine even with smart retries switched on.

A workable default cycle, if you have no data of your own yet: retry at day 1, day 3, day 7, and day 14, restrict access at day 7, and cancel at day 21. Then adjust from what you observe, because recovery rates vary enormously by price point and customer type.

The webhook rules that keep the data honest

Everything above depends on your webhook handler being correct, and this is where generated code is weakest.

  1. Verify the signature. Always. An unverified webhook endpoint is an open API that changes your billing state, and it will be found.

  2. Return 200 fast, process asynchronously. Providers retry on timeout, and a slow handler turns one event into five.

  3. Make handlers idempotent. You will receive duplicate events. Store the provider's event ID and skip anything you have already processed. This is the same pattern as idempotency keys on your own endpoints, applied in the inbound direction.

  4. Handle events arriving out of order. A payment_succeeded can land after a later subscription_updated. Compare timestamps before writing, and ignore events older than your current state.

  5. Never derive state from a single event type. On any billing event, refetch the subscription from the provider's API and reconcile. The event tells you something changed. The API tells you the truth.

Point 5 is the one that eliminates entire categories of bug, and it is worth the extra call.

What the customer actually sees

The technical cycle is half the job. The other half is that a failed payment is usually an accident, and the person wants to fix it.

  • Tell them immediately, on the first failure, with the reason in plain language and a one-click link to update the card. Not a login, not a support ticket, a link.

  • Show it in the product too. Email deliverability is imperfect and a banner in the app reaches people who never opened the message. If your app's mail is unreliable for other reasons, emails going to spam is worth fixing first, because a dunning sequence nobody receives does nothing.

  • Warn before restricting, not at the moment of restricting. "Access will be limited on Friday" recovers payments. "Your access has been limited" recovers complaints.

  • Never delete data at cancellation. Restrict access, keep the data for a defined retention window, and say what that window is. People come back, and the ones who come back to an empty account do not stay.

  • Make the final cancellation explicit and dated. An account that silently stops working is a support ticket you pay for twice.

Test it before it happens to a real customer

Every major provider ships test cards that decline in specific ways, including cards that succeed at checkout and then fail on the first renewal, which is the case you cannot otherwise reproduce without waiting a month. Stripe documents these under testing declines and error codes, and the equivalent exists elsewhere.

Write tests for the sequence, not just the states: first failure, retry, second failure, restriction, recovery on day 12, and the access level at every step. Then test recovery specifically, because the path back to active is the one nobody exercises and the one where customers who just paid you stay locked out.

Seeding this properly is easier with realistic test data in place, since dunning bugs only show up against accounts with history.

Frequently asked questions

Should I use my payment provider's built-in dunning or build my own?

Use theirs for retry scheduling and the card-update emails. Build your own access policy, in-app messaging, and final cancellation rules. The provider cannot know what "restricted" should mean in your product.

How long should a grace period be?

Long enough that an expired card does not cost you a customer, short enough that you are not giving away months. Seven days of full access followed by restriction is a common starting point for monthly consumer subscriptions. Annual and B2B plans usually justify longer, because the recovery value is higher and the buyer is often not the cardholder.

What access should a past-due customer keep?

At minimum, the ability to export their own data and update their payment method. Removing both is how a recoverable billing problem becomes a complaint and a chargeback.

Do I need to handle chargebacks differently from failed payments?

Yes. A chargeback is a dispute with a deadline and evidence requirements, not a retry candidate. Treat it as a separate flow and respond within the provider's window, because unanswered disputes are lost by default.

How do I stop the same bug being reintroduced next time I change the code?

Write the state machine down as a test, not as a comment. A test that walks the full failure and recovery sequence is the only thing that survives a later round of AI-assisted refactoring, which is the same argument as keeping an AI-built app maintainable generally.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.