Dashboard

How to Tell if an AI Vendor's Uptime Claim Is Real

A practical checklist for verifying an AI vendor's uptime claim, from status page history and SLA credits to whether the number covers the model underneath.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
13 September 20261 min read

How to Tell if an AI Vendor's Uptime Claim Is Real

Most "99.9% uptime" badges on an AI vendor's pricing page prove nothing by themselves. A real uptime claim is dated, defined, and checkable against a public record. A marketing one is three nines and a green checkmark with nothing behind it. Before you route a customer-facing feature through someone else's API, it takes about ten minutes to check whether an AI vendor's uptime claim would survive you asking for the receipts. This is a practical checklist for doing exactly that, plus the patterns that should make you slow down before you sign.

What an uptime number actually has to disclose

An uptime percentage with no measurement window attached is not a number, it is a vibe. "99.9% uptime" measured over a rolling 30 days is a real commitment, about 43 minutes of allowed downtime a month. The same figure measured over a full calendar year, with scheduled maintenance quietly excluded from the count, can hide a multi-hour outage without technically breaking the claim. If a vendor states an uptime figure but never says what period it covers or what counts as "down," treat the number as decorative until you can pin it down.

A 5-point checklist for an AI vendor's uptime claim

Run any vendor's reliability claim through these five checks before you build on top of it.

  • A public status page with real history. Not a page that only shows the current state, one with a timestamped incident log going back at least six months, ideally longer.

  • A stated measurement window and exclusions. The claim should say what period it covers and whether scheduled maintenance counts against it. No window means no accountability.

  • SLA credits that actually apply. Read the clause itself. Many SLAs promise a refund only after a formal claim process, capped low, and only for total outages rather than degraded performance.

  • Coverage that includes the model underneath. Plenty of AI products are a thin layer over a foundation model provider. If the wrapper's status page stays green while the underlying model is down, the uptime number is measuring the wrong thing.

  • Evidence outside the vendor's own dashboard. A vendor grading its own homework is not verification. Look for third-party monitoring, public postmortems, or a trial period where you check availability yourself.

Three ways a shaky uptime claim gets dressed up

These are common, generic patterns worth watching for anywhere, not an accusation against any specific company.

  • The badge with no link. A round number like "99.99% uptime" sits on the homepage with no link to a status page, no incident history, and no definition. It is a design element, not a claim you can check.

  • The status page with a short memory. Some status pages only display the last seven or thirty days by default, and older incidents sit somewhere a visitor is unlikely to find. A clean-looking recent history is not the same as a clean track record.

  • The SLA that defines "down" out of existence. An outage clause that only triggers on a total, all-regions failure will rarely fire, even during periods when real requests are timing out or erroring at an elevated rate. Ask what share of requests has to fail, and for how long, before the SLA counts it.

How to test reliability yourself before you commit

A vendor's own claims are a starting point, not a verdict. Run a short test period against the actual endpoint you plan to use, at the concurrency you expect in production, and log latency and error rates yourself. Search the vendor's name alongside terms like "down" or "outage" in developer forums and issue trackers, not just their own changelog. Ask support directly for the postmortems of the last three incidents; a vendor with a real reliability practice will have them written up, and one without will stall. A recent example of the gap between a vendor's self-report and the full picture: when Grok started sending gibberish responses, the vendor called it a rare temporary glitch and moved on, which is exactly the kind of characterization worth checking against your own testing rather than accepting outright. This kind of testing pairs naturally with evaluating a new model release before you switch, since reliability and quality both matter before you move production traffic.

Reliability is one line item, not the whole decision

Uptime is one input among several. Cost matters too, and comparing AI API pricing across providers deserves its own pass rather than an afterthought once you already like a vendor's demo. The same applies to the fine print on limits: a vendor can have excellent uptime and still throttle you into unusable latency under real load, so it helps to understand how AI API rate limits actually work before you commit budget. None of this replaces basic vendor diligence either. The wider set of vendor security and data questions to ask covers what happens to your data, who the subprocessors are, and whether the company behind the product is who the marketing page implies. Building this kind of vendor skepticism into a habit is part of keeping up with AI model releases without burning a day a week on it.

Common questions about AI vendor uptime claims

What counts as a normal uptime number for a production AI API?

Established providers typically publish numbers in the 99.9% to 99.99% range for their core API, measured monthly. Below 99.9% for a service you depend on for revenue is worth a serious conversation with the vendor. Above 99.99% claimed with no public status page to back it up is worth more skepticism, not less.

Does an SLA credit make up for a bad outage?

Rarely in full. Most SLA credits are a percentage of that month's bill, not compensation for lost revenue or damaged trust with your own customers. Treat an SLA as a signal that the vendor takes reliability seriously enough to put a number on paper, not as insurance against the business cost of an outage.

How do I verify a status page isn't cherry-picked?

Scroll back as far as the page allows and look for gaps, not just incidents. A status page with suspiciously few entries over a long period, or one that resets its visible history often, is harder to trust than one with a steady, unremarkable trickle of minor incidents. Real infrastructure has occasional problems. A record with none is either very new or not the whole record.

Should a wrapper's uptime include the underlying model provider's uptime?

Yes, and a vendor that separates the two clearly on its status page is showing you something good about how it operates. If the model provider goes down and requests fail, that is downtime for you regardless of whose infrastructure caused it, and a status page that only tracks the vendor's own servers is not measuring what you actually need to know.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.