How to Roll Out an AI Feature Safely
Shipping a normal feature to one per cent of users answers a clear question: did it break. Shipping an AI feature to one per cent answers nothing unless you decided in advance what "working" looks like, because the output is different every time.
A staged rollout works because you can tell whether the new thing is behaving. Error rate up, roll back. Latency up, roll back. The signal is unambiguous because the code is deterministic: given the same input it does the same thing, and a difference means something changed.
An AI feature gives you none of that. The same input produces different output on Tuesday than on Monday. There is no correct answer to diff against. A five per cent rollout can look completely healthy by every metric your dashboard already tracks while the feature quietly produces bad answers to a specific kind of question you have not thought to test.
So the rollout still happens in stages, in the canary release sense that has been standard practice for years. What changes is what the stages are for.
Decide What "Working" Means Before You Ship
This is the whole job, and it has to happen before any traffic reaches the feature, because after launch you will be reasoning backwards from whatever the numbers happen to say.
Write down, in advance, the numbers that would make you stop. Not aspirations, thresholds, with a decision attached to each:
Signal | What it tells you | A rollback threshold looks like |
|---|---|---|
Fallback rate | How often the model failed, refused or timed out and you served something else | Above 5% sustained over an hour |
Cost per active user | Whether the unit economics survive real behaviour | Above 1.5x your modelled figure |
P95 latency | Whether the feature is usable, not just correct | Above the number you promised the design |
Retry rate | Users re-asking is the clearest signal of a bad answer | Up meaningfully against the control group |
Abandonment | Users starting the feature and leaving | Up against the control group |
Complaints and thumbs-down | The only direct quality signal you will get | Any sustained rise, investigate before continuing |
The middle column matters more than the thresholds. None of these measures correctness, because you cannot measure correctness automatically. They measure whether users behave like people being helped. A user who asks the same question three times and leaves has told you the answer was wrong without ever clicking a rating.
Build the Off Switch First
Before the feature ships to anyone, you need to be able to turn it off without a deploy. Not a config change that requires a restart, not a branch revert: a flag you can flip in seconds while looking at a graph.
There are two switches, and they are different:
The rollout flag, controlling what percentage of users see the feature. This is how you go from 1% to 100%, and back.
The kill switch, which disables the AI path entirely and falls back to whatever the product did before. This is what you use at three in the morning when something is badly wrong and you do not yet know what.
The kill switch needs somewhere to fall back to, which is the part teams skip. If the AI path is the only path, turning it off breaks the feature rather than degrading it. Decide early whether the fallback is the old non-AI behaviour, a cheaper model, a cached response or an honest message, and build it at the same time as the feature rather than afterwards. If you do not have flags in place at all, adding feature flags to an AI-built app is the prerequisite for everything here.
The Stages, and What Each One Is Actually For
Internal only
Your team, with no percentage attached. The goal is not statistics, it is to have people who understand the system read a few hundred real outputs. This is the only stage where anyone will look at quality directly, so make it count: read the outputs for the queries you did not anticipate, not the ones you built the demo around.
One per cent
Now you are testing the plumbing, not the quality. Does it stay up. Does latency hold under real concurrency. Does the fallback path actually trigger when the provider returns a 500, which you should verify by forcing one rather than waiting to find out. Is the cost per user anywhere near your model.
One per cent of a small product is too few users to tell you anything about quality, and that is fine, because quality is not what this stage is for.
Ten per cent, and a control group
This is the first stage that can tell you something real, and only if you hold back a comparison group. Retry rate and abandonment mean nothing in isolation; they mean something against users who did not get the feature. Run it long enough to cross a weekend, because weekday and weekend usage are different products.
Fifty, then a hundred
By now the behavioural signals are stable and you are mostly watching cost and capacity. This is where rate limits get tested for real, and where a provider quota you never came close to in testing becomes the thing that pages you.
What to Watch That Is Not on Your Dashboard
Cost per user by percentile, not the average. A heavy tail is normal, and an average that looks fine can hide users costing fifty times the median. The full calculation is in working out what an AI feature costs per user.
The inputs that produce fallbacks. These cluster, and the cluster is usually a category of user need the feature handles badly. This is the highest-value thing you will learn during a rollout.
Output length drift. A model quietly returning longer responses moves your costs and your latency at once, and nothing else on your dashboard will explain it.
Provider incidents. Your feature's availability now includes someone else's. Subscribe to their status page before you need it.
Log the prompt, the model, the response and the outcome for every call from the first day. Sampling is tempting and you will regret it the first time you need to understand a complaint from a specific user on a specific afternoon.
Rolling Back Is Not a Failure
The reason to define thresholds in advance is that in the moment, with a feature you spent three weeks building, every bad number has a plausible explanation. Traffic was unusual. That cohort is atypical. It will settle. Sometimes that is even true, and it is not a judgement you can make fairly while looking at your own work.
A threshold written down a week earlier, by you, in a calmer state, is the only defence against that. Hit it, roll back, investigate, ship again. A feature that went out and came back twice before working is an ordinary outcome, and it costs far less than one that stayed out while the evidence accumulated.
Plan the other end too. Features get retired, and an AI feature with a running cost is retired more often than most, which is the subject of sunsetting an AI feature. Knowing what the off-ramp looks like makes the decision to build easier, not harder.
For the broader sequence of getting something built and into users' hands, how to build an app with AI covers the ground before this point. The rollout is where an AI feature stops being a demo, and it is worth treating as its own piece of work rather than the last hour of the last day.
Frequently Asked Questions
Why can I not roll out an AI feature like any other feature?
Because there is no correct output to compare against. Standard rollout signals detect breakage, and an AI feature can produce bad answers while every technical metric looks healthy.
What metrics should I watch during an AI feature rollout?
Fallback rate, cost per active user by percentile, p95 latency, retry rate, abandonment and direct complaints. The behavioural ones need a control group to be meaningful.
What percentage should I start at?
Internal first, then one per cent to test the plumbing, then ten per cent with a held-back control group, which is the first stage that tells you anything about quality. The exact numbers matter less than holding back a comparison group.
Do I need a kill switch separate from the rollout flag?
Yes. The rollout flag controls exposure; the kill switch disables the AI path entirely and falls back to something else. You need the second one at three in the morning when you do not yet know what is wrong.
How long should each stage run?
Long enough to cross a weekend at minimum, since weekday and weekend usage differ. Beyond that, until the behavioural signals stop moving rather than until a fixed date passes.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


