Why AI Image Models Cannot Spell
Ask an image model for a shop sign reading OPEN and you may get OPFN, OPFM, or something that looks like a word in a language nobody speaks. The reason is structural, and understanding it tells you when to stop fighting the prompt.
An image model garbles text because it was never handling text. It is painting the visual texture of writing, letter-like shapes in a letter-like arrangement, the way it paints the texture of brick or fur. Brick that is slightly wrong still reads as brick. A word that is slightly wrong reads as a mistake, because writing is the one thing in a photograph where a human eye checks every single element against a known answer.
That asymmetry, not a lack of capability, is why this failure is so conspicuous.
The Model Never Sees Letters
A language model breaks your words into tokens and works with them as discrete symbols. An image model does something different: your prompt is encoded into a numerical description of meaning, and that description conditions a denoising process that shapes pixels. The pipeline is explained in more detail in how diffusion models generate images, but the step that matters here is that nothing downstream of the encoder is holding a list of characters.
So when you ask for a sign reading BAKERY, the model does not receive an instruction to place B, then A, then K. It receives a region of meaning that includes "sign", "text", "bakery-ish", and it generates pixels consistent with that. Six letter-shaped forms in a row is consistent with that. So is five. So is BAKFRY.
The model is not spelling badly. It is not spelling.
Why Some Letters Come Out Right
Common words that appear constantly in training data, OPEN, STOP, CAFE, SALE, have been seen so often as a specific visual pattern that the model has effectively memorised their shape. They behave less like spelled words and more like logos. That is why short, ubiquitous words work far more reliably than long or unusual ones, and why a made-up brand name is the hardest possible request.
Request | Typical outcome | Why |
|---|---|---|
A sign reading OPEN | Usually correct | Memorised as a visual unit, seen constantly |
A sign reading CLOSED | Often correct | Common, but long enough to drift |
A poster reading "Tuesday Night Quiz" | Partly wrong | Multiple words, no memorised composite form |
A logo reading "Vantrell" | Usually wrong | Invented word with no training precedent |
A paragraph of body text | Reliably gibberish | Too many characters, no per-character control |
The pattern is consistent: the further a string is from something the model has seen rendered thousands of times, the more it falls back on producing writing-shaped texture.
Why Newer Models Got Better at It
Text rendering has improved noticeably, and it is worth being precise about why, because the improvement does not change the underlying nature of the system. Three things drove it:
Character-aware text encoders. Google researchers showed in Character-Aware Models Improve Visual Text Rendering that giving the encoder character-level features produces large accuracy gains, and the biggest gains land on exactly the rare words that fail worst.
Training data that includes far more rendered text with known ground truth, which turns spelling from an accident into something the model is explicitly rewarded for.
Higher native resolution. Legible small text needs pixels. A model generating at 2048 by 2048 has room to form clean letterforms that a 512 by 512 model physically cannot resolve.
All three shift the odds. None of them gives the model a character buffer it can check its work against. This is why even strong models still fail on long strings: the failure mode is unchanged, it just starts later.
How to Work With It
Once you accept that the model approximates writing rather than composing it, the workarounds are obvious:
Keep requested text short. One or two common words succeed far more often than a phrase.
Ask for text as a subject, not as a detail. "A neon sign reading OPEN, centred, filling the frame" gives the model resolution to work with. The same words rendered small in a background will smear.
Generate the image without text and add the typography yourself. For anything that has to be correct, a logo, a price, a product name, this is the reliable answer rather than a fallback.
Use an editing model to replace the text in a finished image, which is a far more constrained task than generating it from scratch.
Generate several and pick. Spelling is stochastic, so a batch of six will usually contain one that landed.
The prompt-side techniques in writing better AI image generation prompts help with framing and style, and they will improve your odds on text, but no prompt converts an approximation into a guarantee. If the words must be right, put them there yourself.
The Same Root as the Counting Problem
This belongs to a family of failures where a model produces something statistically plausible instead of something verified. Why AI models are bad at counting describes the language-model version of the same gap: the output looks like the shape of a correct answer without any step that checked it.
Recognising the family is useful beyond image generation. Any time a model is producing output whose correctness a human can verify instantly but the model itself never checked, expect confident near-misses. That is a general property of how these systems work, covered further in how AI models actually work, and it is the right prior to bring to any new model's claims.
Frequently Asked Questions
Why does AI generate gibberish text in images?
Because it generates the visual appearance of writing rather than placing specific characters. Letter-like shapes in a plausible arrangement satisfy what the model is optimising for, and nothing in the pipeline checks the result against the word you asked for.
Which AI image generators are best at text?
The ones with stronger text encoders and higher native resolution do noticeably better, and this changes with every release. Rather than tracking the leaderboard, test your specific strings: the same model can nail a one-word sign and mangle a three-word headline.
Will AI image models ever spell reliably?
They are improving steadily, and short common text is already close to dependable. Long or invented strings remain unreliable because the architecture approximates rather than composes. Expect the failure threshold to move, not disappear.
How do I get correct text in an AI-generated image?
Generate the image without the text and add it in a design tool, or use an editing model to replace the text in a finished image. For anything that must be exactly right, do not leave it to generation.
Why does it get short words like OPEN right?
Very common words appear so often in training data that the model has effectively memorised them as visual units, closer to a logo than a spelled word.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


