Dashboard

Staging and Production for an AI-Built App

Most AI-built apps ship straight to production because a second environment sounds like infrastructure work. It is an afternoon, and it is mostly config.

Steve Jefferson
Steve Jefferson
Developer Advocate
31 August 20261 min read

Most apps built with AI go straight from a laptop to production, because a second environment sounds like infrastructure work and the tutorials skip it. It is not infrastructure work. It is one more deployment of code you already have, pointed at a different database, and it takes an afternoon.

What makes it useful is being precise about the differences. A staging environment that differs from production in the wrong ways is worse than none, because it produces confident green results that mean nothing.

The four things that must differ

  1. The database. A separate instance, not a separate schema in the same one. Shared instances mean a bad migration on staging takes production with it, which is the exact event you built staging to prevent.

  2. Every external credential. Test keys for the payment provider, a sandbox mail domain, a separate storage bucket. See environment variables and secrets for how to keep the two sets apart.

  3. Outbound side effects. Staging must not send email to real addresses, charge real cards, or post to real webhooks. Route mail to a catcher and point webhooks at a request bin.

  4. Search engine visibility. A noindex header and a robots rule on the staging host. Duplicate staging content in search results is a genuinely annoying problem to unpick later.

The three things that must stay identical

  • The runtime version. Same language version, same major dependency versions, ideally the same container image. A staging pass on a different runtime is not evidence.

  • The migration path. Staging gets migrations applied the same way production will, in the same order, from the same files. Hand-editing a staging schema to fix something destroys the only signal it produces.

  • The build. Deploy the same artefact you will promote. Rebuilding for production means the thing you tested is not the thing you shipped.

That last one is the difference between a staging environment and a demo site. If the artefact changes between the two, you have tested a sibling of your release.

The seed data trap

This is where most staging environments quietly stop being useful. They get seeded once with a dozen tidy records, and every test after that runs against data that looks nothing like production: no unicode names, no customer with 4,000 rows, no half-finished signup from 2024, no record with a null in the column your code assumes is always set.

Three workable options, in order of effort:

Approach

Realism

Effort

Main risk

Hand-written seed script

Low

An hour

Only tests the happy path you imagined

Generated volume data

Medium

Half a day

Realistic size, unrealistic shape

Anonymised production dump

High

A day, plus ongoing

Anonymisation gaps leak real personal data

Start with the seed script and add the awkward cases as you find them. Every production bug that surprises you should leave behind one new seed row, which is how the script becomes genuinely useful over a few months rather than by planning.

If you take the anonymised dump route, understand that you are now storing personal data in a second place, with the obligations that come with it. Partial anonymisation is common and is where the incidents happen.

The difference is configuration, not code

If the two environments differ anywhere in your source, you have built two applications. Every difference belongs in configuration read at startup, which is the dev and production parity principle and the reason the same artefact can be promoted rather than rebuilt.

python
# wrong: the environment is baked into the code
if ENV == "staging":
    mailer = ConsoleMailer()
else:
    mailer = SendGridMailer(api_key="SG.live...")

# right: the environment is a value the code reads
mailer = build_mailer(
    driver=env("MAIL_DRIVER"),       # 'log' on staging, 'sendgrid' in production
    api_key=env("MAIL_API_KEY"),
)

The first version has a staging-only branch that never runs in production and is therefore never tested there. Conditionals on environment name are how a bug ships that only exists in the environment you cannot rehearse.

A useful check: search your codebase for the string production. Every hit is a place where the environments diverge in code rather than config, and each one is a small hole in the value staging provides. Environment variables and secrets covers keeping the two sets of values apart safely.

Do you need a third environment?

Almost certainly not. A separate development, staging and production trio is standard in larger teams because many people need to work without colliding. Solo, the local machine is the development environment and a third deployed environment is a thing to keep in sync for no return.

The exception is a customer who needs somewhere to test against. That is a demo or sandbox environment, it has different requirements from staging, and it should not be your staging environment, because you will stop being willing to break it.

Promoting a change

The whole point is a repeatable sequence that a tired person can follow at nine on a Friday evening:

  1. Merge to the staging branch. Deploy to staging automatically.

  2. Run migrations on staging. If they fail, the release stops here and costs nothing.

  3. Walk the two or three flows that make you money. Signup, checkout, and whatever the app is actually for.

  4. Promote the exact artefact to production. Not a rebuild.

  5. Run migrations on production, then walk the same flows again.

Steps three and five being the same walk is deliberate. A discrepancy between them is the most informative signal the whole setup produces, and it is what tells you an environment difference has crept in.

What this costs

On most managed platforms, a staging environment is a second small instance and a second small database, which is usually a low double-digit monthly figure and sometimes free at hobby scale. Check the pricing page for your own platform rather than trusting a number in an article, since these change often.

Set the staging instance to the smallest tier available. It is not serving traffic, and a staging environment sized like production is the most common way this becomes expensive enough to switch off.

Letting an agent use it

A staging environment is also what makes it reasonable to let a coding agent run things that touch a database, since the blast radius is a machine you can rebuild. That has its own rules, covered in giving a coding agent staging access.

Rehearsing the risky migrations

The single highest value use of staging is running destructive migrations against realistic data before they touch anything you cannot rebuild. Four to check every time:

  • Dropping or renaming a column. Confirm nothing still reads it, including the reports and exports nobody remembers writing.

  • Adding a not-null constraint. On an empty staging table this passes instantly and on production data it fails, so this is precisely the case that needs realistic rows.

  • Adding an index to a large table. Time it. A migration that takes eleven minutes on production may hold a lock for eleven minutes.

  • Any data transformation. Run it, then count the rows it changed and check a sample by hand before you trust it.

Time each migration on staging and write the number down. A deploy plan that says the schema change takes about four minutes is a very different conversation from one that says it should be quick, and it is the difference between deploying an app built with AI calmly and doing it hopefully. The rest of the pre-launch checks are in testing before launch.

Frequently asked questions

Do I need staging if I am the only user so far?

Not on day one. The moment a person who is not you depends on the app, or the moment you have data you would be upset to lose, the answer changes. Migrations are the usual trigger: the first destructive migration you run without a rehearsal is the one that teaches this lesson expensively.

Is a preview deployment per branch the same thing?

Close, and better in some ways, provided each preview gets its own database and its own test credentials. Preview environments that share one database with each other reproduce the shared-instance problem with extra steps.

How do I keep staging from drifting out of date?

Deploy to it on every merge, automatically, whether or not you plan to test. A staging environment that is updated only when someone remembers is three weeks stale exactly when you need it, and when the app breaks in production is a worse place to discover that.

Should staging use the same AI model and provider?

Same model, separate API key with its own spending limit. Different models behave differently enough that a staging pass on a cheaper model tells you very little about production behaviour, and the separate key means a runaway loop in testing cannot exhaust the budget your live app depends on.

This staging/production split matters just as much for internal tools as it does for customer-facing apps. See our worked example on turning a Notion doc into an app with AI, which walks through building a real internal tool from a client/project database, including getting per-client permissions right before anyone but you touches it.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.