How to Add Maintenance Mode to an AI-Built App
Maintenance mode needs a 503, a Retry-After header and a way to let yourself in. Here is how to add it to an AI-built app without locking yourself out.
Ask an agent to add maintenance mode to an AI-built app and you will get a boolean, a check in your routing layer and a page that says "We will be back soon." It works, it takes four minutes, and it is missing the two things that make maintenance mode safe to use: a way for you to stay inside the app while it is on, and the right HTTP status so that an hour of planned downtime does not read to a search engine like a site that has been deleted.
This walks through all four parts. The examples use Express and nginx because that is where most AI-built app stacks land, but the shape transfers to anything.
1. Put the switch somewhere you can reach when the app is broken
The first mistake is storing the flag in the database. If you are taking the app down to run a migration, the database is the thing that might be unavailable, and now you cannot turn maintenance mode off from the admin panel because the admin panel cannot read the flag.
Put it somewhere with a different failure domain from the thing you are maintaining. In rough order of preference:
An environment variable, if your host can redeploy config without a full build. Simple and completely independent of your data layer.
A key in a cache or key-value store that is not the primary database.
A file on disk, for a single-server deployment.
The pattern that reliably fails is a settings row in the same Postgres instance you are migrating.
2. Write the middleware, with the allowlist
The middleware itself is short. The allowlist is the part that matters, because maintenance mode is useless if it also locks out the person doing the maintenance. You need to verify the migration worked before you let users back in.
const MAINTENANCE = process.env.MAINTENANCE_MODE === 'on';
const BYPASS = process.env.MAINTENANCE_BYPASS_TOKEN;
// Paths that must keep working while the app is down.
const ALWAYS_OPEN = ['/healthz', '/maintenance', '/webhooks/stripe'];
function maintenance(req, res, next) {
if (!MAINTENANCE) return next();
if (ALWAYS_OPEN.some((p) => req.path.startsWith(p))) return next();
// Anyone holding the bypass cookie gets the real app.
if (BYPASS && req.cookies?.maint_bypass === BYPASS) return next();
// Set the cookie once via ?bypass=<token>, then browse normally.
if (BYPASS && req.query.bypass === BYPASS) {
res.cookie('maint_bypass', BYPASS, { httpOnly: true, sameSite: 'lax' });
return res.redirect(req.path);
}
res.set('Retry-After', '900');
res.status(503);
return req.accepts('html')
? res.render('maintenance')
: res.json({ error: 'maintenance', retry_after: 900 });
}
app.use(maintenance);Three details in there are worth stating explicitly, because an agent asked for "maintenance mode" will not add them unprompted.
The health check stays open. If your platform probes /healthz and you return 503 to it, the platform will conclude the instance is unhealthy and start restarting or replacing containers while you are trying to work on the database. Keep the probe path out of maintenance mode, or point the probe at something that reports process health rather than application readiness.
Inbound webhooks need a decision. A payment provider posting to /webhooks/stripe during your maintenance window will get a 503. Good providers retry, so that is often acceptable, but only if you have checked that yours does and that its retry window is longer than your outage. If the webhook handler only writes to a queue, leave it open. If it writes to the database you are migrating, let it 503 and rely on retries, deliberately rather than by accident.
The bypass is a query parameter once, then a cookie. Query-string-only bypass breaks as soon as you click a link, and you will spend the outage appending the token to every URL.
3. Return 503, not 200 and not 302
This is the part with consequences that outlast the maintenance window.
A maintenance page served with a 200 tells every crawler that your pages now legitimately contain the text "We will be back soon." A redirect to /maintenance tells them your content has moved there. Both can cost you indexed pages over a long outage.
A 503 says the server is temporarily unable to handle the request, and the Retry-After header tells the client how long to wait. Google Search Central's guidance on planned downtime is explicit that a 503 with Retry-After is the correct signal and that short outages handled this way do not harm a site.
Set Retry-After to slightly more than your honest estimate. Underestimating and being asked again immediately is worse than overestimating.
4. Put the switch above the app when you can
Everything above assumes your application process is running. Sometimes it is not, because the thing you are doing is replacing it.
If you have a reverse proxy or CDN in front, that is a better place for the switch, because it works even when nothing is listening behind it:
server {
# ...
if (-f /etc/nginx/maintenance.on) {
return 503;
}
error_page 503 @maintenance;
location @maintenance {
root /var/www/maintenance;
rewrite ^(.*)$ /index.html break;
add_header Retry-After 900 always;
}
}Now maintenance mode is `touch /etc/nginx/maintenance.on` and a reload, and it holds while the application containers are stopped entirely. The application-level middleware is still worth keeping for the allowlist, since the proxy has a harder time distinguishing you from everyone else.
What to tell the agent
If you are having an AI coding agent build this, the prompt that produces the right thing on the first pass names the parts rather than the feature:
Add maintenance mode. Read the flag from an environment variable, not the database. Return 503 with a Retry-After header, never 200 or a redirect. Keep /healthz and inbound webhook paths exempt. Support a bypass token passed as a query parameter that sets an httpOnly cookie, so an admin can use the app normally while maintenance mode is on. Serve JSON to API clients and HTML to browsers.
That is the general lesson about asking agents for named features: the name maps to the obvious 80%, and the remaining 20% is where the outage happens. The same pattern shows up in rolling out an AI feature safely.
Test it before you need it
Maintenance mode is code that runs once a quarter, under pressure, at the exact moment you cannot afford it to be wrong. That is the profile of code that is always broken when it is finally used.
Test it in daylight, on a Tuesday, when nothing is wrong:
Turn it on in staging and load the app in a normal browser. You should see the page.
Check the status code, not just the page. `curl -I https://staging.example.com/` should print `HTTP/1.1 503` and a `Retry-After` header. A page that looks right with a 200 behind it is the failure this whole article is about.
Hit the bypass URL, then browse normally. You should reach the real app without re-appending the token.
Curl your health check path. It must return 200 while maintenance mode is on.
Request an API path with `Accept: application/json`. You should get JSON with a 503, not an HTML page a client cannot parse.
Turn it off and confirm everything returns to normal without a restart.
Six minutes, once, and the version you run during an actual migration is a version you have seen work.
Maintenance mode is a blunt instrument
Before you reach for it, check whether you need it. Most changes people take an app down for do not require it. Additive migrations, anything behind a feature flag, and most deploys can happen live. Reserve maintenance mode for destructive migrations, data moves between systems, and the cases where serving stale or partial data would be worse than serving nothing.
The general rule is that the window should be as short as the riskiest single operation inside it, not as long as the whole change. If a two hour job has ten minutes of destructive work in the middle, take the app down for the ten minutes.
When you do use it, tell people where to look. A maintenance page that links to a status page converts a support ticket into a page view. Keeping this kind of operational scaffolding in shape is most of what ongoing maintenance of an AI-built app actually consists of, and it is worth building before the first unplanned outage rather than during it.
FAQ
What status code should a maintenance page return?
503 Service Unavailable, with a Retry-After header. A 200 tells crawlers the maintenance text is your real content, and a redirect tells them your content moved.
How do I avoid locking myself out of my own app?
Use a bypass token that sets a cookie, checked before the maintenance response. Store the maintenance flag somewhere independent of the database, so you can turn it off even when the database is down.
Will maintenance mode hurt my search rankings?
Not for a short, correctly signalled outage. Google's guidance is that a 503 with Retry-After is the right way to handle planned downtime. Risk rises with duration, and with serving 200s or redirects instead of 503s.
Should webhook endpoints stay available during maintenance?
It depends on what they touch. If the handler only enqueues work, keep it open. If it writes to the database you are migrating, let it return 503 and rely on the provider's retries, after confirming the retry window is longer than your outage.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


