Dashboard

How to Prompt AI to Write an Incident Update

The hard part is publishing something useful at minute twelve, before you know the cause. Here is the prompt constraint that stops a model guessing.

Steve Jefferson
Steve Jefferson
Developer Advocate
14 September 20261 min read

How to Prompt AI to Write an Incident Update

The hard part of an incident update is not the writing. It is that you have to publish something useful at minute twelve, when you do not yet know the cause, and every instinct pushes you either into silence or into a guess you will have to retract. A model will happily supply the guess. Prompted correctly it will do the opposite, and that constraint is the whole technique.

The reliable way to prompt AI to write an incident update is to hand it facts and forbid it from adding any. What follows is one prompt that produces three drafts at three points in an incident, plus the specific instruction that stops the model naming a cause you have not established.

The rule that makes this work

Give the model facts and forbid inference. Language models are pattern completers, and the pattern that follows "customers are seeing 500 errors on checkout" in the training data is a sentence about a database or a deploy. Left alone, the model will write that sentence with total confidence, and you will publish a cause you invented.

Every sentence in the update must be traceable to something on this list. If a claim is not on the list, write that it is unknown.

That line, pasted verbatim into the prompt, does more than any amount of tone guidance. It converts the model from an author into a formatter, which is what you actually want while an incident is live.

The prompt to write an incident update

Fill in the facts block with whatever you genuinely know, including the things you know you do not know. The shorter the list, the shorter the update, which is correct.

text
You are drafting a public status update for customers during a live incident.

FACTS (the only things known to be true):
- Started: 09:14 UTC, detected by automated alerting
- Symptom: checkout returns an error for some customers
- Scope: affects card payments only; the rest of the app is responding normally
- Estimated share affected: unknown
- Cause: unknown, under investigation
- Workaround: none available
- Data loss: none observed, not yet confirmed
- Next update: within 30 minutes

RULES:
- Every sentence must be traceable to the FACTS list. If something is not on the
  list, say it is unknown. Never infer or suggest a cause.
- No apology adjectives (deeply, sincerely, extremely). One plain apology sentence.
- No phrases like 'a small number of users' unless a number is in FACTS.
- Lead with what the customer cannot do right now.
- Plain sentences, no marketing tone, 90 words maximum.

Write the update.

Two of those rules earn their place repeatedly. Banning "a small number of users" matters because the model reaches for it automatically and it is the single most distrusted phrase in status page history. Banning apology adverbs matters because they scale inversely with credibility: the more deeply a company is sorry at minute twelve, the less the reader believes the rest of the sentence.

How to prompt AI to write an incident update at three points

Reuse the same prompt at three points, changing only the FACTS block and one line of instruction. The drafts should not read the same, because they are doing different work.

The first update, inside fifteen minutes

The job is acknowledgement, not explanation. The reader wants to know that you know, and what they cannot currently do. Cause is almost always unknown and should be stated as unknown. Add this line to RULES:

text
- This is the first update. Do not speculate on cause or timeline. State what is
  affected, what is not, and when the next update will come.

The middle update, while you are still working

This is the one people skip, and skipping it is what turns an incident into a reputational event. The job is to prove the clock is still running. Even an update that says nothing new has value if it is on time. Add:

text
- This is a progress update. It is acceptable to report no change. Say what has been
  ruled out, if anything. Do not restate the whole incident, reference it and add
  only what is new since the last update.

Reporting what you have ruled out is the trick here. "We have confirmed this is not a payment provider outage" is real information, costs you nothing, and reads as progress even when you have not fixed anything.

The resolution update

Now, and only now, the cause can appear, because now you know it. This update carries the recovery time, the confirmed scope, and whether anything needs the customer to act. Add:

text
- This is the resolution update. Include: what was wrong, when service returned to
  normal, the confirmed scope of impact, and any action the customer needs to take.
- Do not promise that it will never happen again. Describe only changes already made
  or already scheduled.

The last rule prevents the most common own goal in incident writing. A resolution note that promises permanence gets quoted back at you the next time. Describing the change you actually made is both more honest and more reassuring. The deeper analysis belongs somewhere else entirely, which is what a separate postmortem is for.

What to check before you publish

Read the draft once against four questions. This takes about twenty seconds and catches nearly everything.

  1. Is there a cause in here that I have not confirmed? Delete it.

  2. Does the first sentence say what the customer cannot do? If it opens with an apology or with the word "currently", rewrite the first sentence.

  3. Is there a number I cannot defend? Percentages and user counts invented by a model are the fastest way to lose the thread.

  4. Have I committed to a next update time I will actually hit? If unsure, widen it. Missing your own stated deadline costs more than a longer interval would have.

None of this is a substitute for a process. If you want the underlying discipline rather than the writing layer, Google's SRE book chapter on managing incidents remains the clearest public description of how roles and communication split during an outage, and it is where the habit of a dedicated communications lead comes from.

Where this fits with your other incident writing

An incident update is not an apology email and it is not a postmortem, and using one where another belongs is the usual failure. The update runs during the event and states facts. The apology email comes afterwards, to affected customers specifically, and carries remedy rather than status. The postmortem is internal and carries cause.

If you publish updates regularly, they need somewhere to live that is not your main app, since the main app is the thing that is down. That is the argument for a status page hosted separately from the product. And if the incident is not yours at all but your model provider's, the degradation decision comes before the writing does.

The same raw-material-in approach works for looking back across a whole sprint rather than one event: see how to prompt AI to write a sprint retrospective.

Frequently asked questions

Should I let AI publish updates automatically during an incident?

No. Draft with it, publish yourself. The failure mode is not bad prose, it is a confident sentence about a cause nobody confirmed, and a human reading the draft catches that in seconds.

How often should the updates go out?

Pick an interval and state it in every update, then hold it. Thirty minutes suits most small products. The interval matters more than the content, because a missed update reads as a worse incident than a boring one.

What if I genuinely have nothing new to say?

Say that. "No change since 10:45. Still investigating, next update at 11:45." Readers handle a slow fix far better than a silent one, and the prompt above is built to allow an update that reports no progress.

Can I reuse one prompt across different products?

Yes, and you should, because consistency during an incident is a feature. Keep the RULES block fixed and swap only the FACTS. That is the same reasoning behind building a reusable prompt library rather than rewriting from scratch under pressure.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.