Dashboard

How to Prompt AI to Write Alt Text

One describe-the-image prompt produces alt text that fails accessibility review. Branch the prompt on the image's function instead.

Steve Jefferson
Steve Jefferson
Developer Advocate
31 August 20261 min read

The single most common alt text prompt is describe this image, and it produces text that fails review roughly three times out of four. Not because the description is wrong. Because alt text is not a description. It is a replacement, and what it should say depends entirely on why the image is on the page.

A photo of a padlock next to a paragraph about encryption needs no alt text at all. The same photo used as a link to a security page needs alt text describing the destination, not the padlock. Same pixels, opposite answers.

Decide the role before you write the prompt

The W3C alt decision tree sorts every image into a handful of roles. Four of them cover almost everything you will meet:

Role

Example

What the alt must do

Informative

A chart in an article, a product photo

Convey the information the image carries

Decorative

A stock photo beside a paragraph that already says it

Be empty, so screen readers skip it

Functional

A logo that links home, an icon button

Describe the action or destination, not the picture

Complex

A dense diagram or data visualisation

Give a short label plus a longer description nearby

Getting the role right is most of the work. A model cannot infer it from the image alone, because the role lives in the page, not the pixels. You have to tell it.

Informative images

Give the model the surrounding context and the length limit, and tell it what to leave out. Redundancy with nearby text is the most common failure here, and it is genuinely annoying to listen to.

text
Write alt text for this image.

Role: informative. It carries information not stated in the text.
Surrounding paragraph: "Support tickets rose through Q2 before
flattening in July after the self-serve refund flow shipped."
Limit: 125 characters.

Rules:
- Do not start with "image of" or "photo of".
- Do not repeat facts already in the surrounding paragraph.
- State what a sighted reader would take from it that the text omits.
- Plain sentence, no trailing full stop needed.

Return only the alt text.

For a chart, the useful answer is the shape of the data, not the chart furniture. Ticket volume climbing steeply from April to June then flat through July beats bar chart showing support tickets by month.

Decorative images

The correct alt text is an empty string, and this is where models fight you hardest. Asked to describe a decorative image, a model will always produce something, because producing nothing feels like failing the task.

So do not ask it to describe. Ask it to classify, and handle the empty case in your own code:

text
Classify this image's role on the page. Context: it sits directly
below the heading "How refunds work" and above a paragraph that
fully explains the refund process. The image is a stock photo of
a person at a laptop.

Answer with exactly one of: INFORMATIVE, DECORATIVE, FUNCTIONAL, COMPLEX.
Then, on a new line, one sentence of justification.
If DECORATIVE, the alt attribute must be empty.

Then in your template, alt="" for the decorative case. Not "decorative image", not a space, not the filename. An empty alt attribute is what tells a screen reader to skip the element entirely; a missing alt attribute makes it read the filename out loud, which is the worst outcome available.

Functional images

When the image is inside a link or a button, the alt text is the label for that control. Describing the artwork here is actively harmful, because the user needs to know where they will land.

text
Write the alt text for an image used as a control.

The image: a magnifying glass icon.
It is: a button that opens site search.
Limit: 60 characters.

The alt text must name the action or destination, not the picture.
Return only the alt text.

Correct output: Search. Not magnifying glass icon, and not search icon, which describes the picture rather than the action. This distinction accounts for a large share of real accessibility findings on otherwise careful sites, and it is covered further in making an AI-built app accessible.

Complex images

A dense diagram will not fit in 125 characters, and cramming it in serves nobody. Split it: a short alt attribute that identifies the image, and a full description in the page near it, available to everyone.

text
This is a complex image. Produce two outputs.

1. ALT: under 100 characters. Identify what the diagram is and
   point to the full description. Do not attempt to summarise the data.
2. LONGDESC: a structured description for a caption or details
   element. Lead with the overall shape or conclusion, then walk
   the components in reading order. Under 200 words.

Label each output clearly.

One prompt that handles the branching

Four prompts is right when you are working image by image. For a batch, fold the classification and the writing into a single call so the model commits to a role before it writes, rather than defaulting to description:

text
For the attached image, do this in order.

STEP 1. Using the page context below, classify the role:
INFORMATIVE / DECORATIVE / FUNCTIONAL / COMPLEX.

STEP 2. Write the alt attribute according to the role:
- DECORATIVE -> output exactly: ""
- FUNCTIONAL -> the action or destination, under 60 characters
- INFORMATIVE -> the information the image adds, under 125 characters
- COMPLEX -> a short identifier plus "described below", then a
  separate LONGDESC block under 200 words

STEP 3. State in one line why the role is not one of the other three.

Page context: <heading, the paragraph before, the paragraph after,
and whether the image is wrapped in a link or a button>

Output as: ROLE / ALT / WHY-NOT / LONGDESC (omit LONGDESC if not COMPLEX).

Step 3 is what earns its place. Forcing the model to argue against the other roles catches the decorative case, which it otherwise almost never chooses, and it gives you a line to skim when you review a hundred of these.

Feeding the surrounding paragraphs rather than the whole page keeps the context useful. This is the same discipline as prompting with a screenshot: what you crop out matters as much as what you send.

The failures worth knowing about

Failure

What it looks like

Why it happens

Reading invented text

Alt text quoting words that are not in the image

Models complete plausible signage and labels

Naming brands

A laptop becomes a specific manufacturer's model

Visual similarity plus a confident prior

Restating the caption

Alt text and caption read identically aloud

The caption was in the context and got echoed

Refusing to be empty

decorative image or spacer as alt text

Producing nothing reads as failing the task

Describing composition

close-up shot with shallow depth of field

Training on photography captions rather than accessibility text

The first two put false statements in front of people who cannot check them, which makes them different in kind from the others. If the alt text asserts a specific word, number or brand, verify it or cut it.

The last row is worth a template rule rather than a per-image fix. If your pipeline fills a field, filling a template without breaking it covers keeping an empty value genuinely empty rather than helpfully populated.

Reviewing what comes back

Four checks, worth about thirty seconds per image:

  1. Read it in place of the image. Does the paragraph still make sense, or does it now say the same thing twice?

  2. Count characters. Over about 125 and most screen readers keep going anyway, but the length usually signals that the role was misjudged.

  3. Check for hallucinated specifics. Models invent readable text in photographs, brand names on products, and numbers on charts. If the alt text quotes a value, verify it against the image.

  4. Check nothing begins with image of, picture of, or graphic showing. Screen readers already announce that it is an image.

The hallucination check is the one to keep. It is the only failure mode here that puts a confident false statement in front of a user who has no way to notice.

Frequently asked questions

How long should alt text be?

There is no limit in the specification. The 125 character convention comes from older screen reader behaviour and remains a good discipline for informative images. Complex images are the exception and should use a short alt plus a longer description, per the WCAG guidance on non-text content.

Does alt text help SEO?

Mildly, and it is the wrong reason to write it. Alt text stuffed with keywords reads terribly aloud and is the sort of thing search engines have discounted for years. Write it for the listener and take the ranking benefit as a side effect. SEO for an app with no marketing budget covers where the real gains are.

Can I run this over hundreds of images at once?

Yes, if you pass the page context per image rather than the image alone, and if you review the decorative classifications by hand. Batch jobs tend to mark almost nothing decorative, which produces a site where a screen reader narrates every background photo.

What about images inside an image generation workflow?

The generation prompt is not the alt text, and reusing it produces alt text full of style directions like cinematic lighting. Write the alt text from the finished image. The two jobs are covered separately in writing image generation prompts.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.