How to Verify an AI Translation You Cannot Read
You can ship a Spanish landing page without speaking Spanish, but not by trusting the model and hoping. Here is a verification loop that catches meaning errors, register mistakes, and silently dropped content before a customer finds them.
You cannot proofread a language you do not speak, but you can verify an AI translation without reading a word of it. The trick is to stop asking "is this good Spanish?" and start asking questions that have checkable answers: does the round trip come back saying the same thing, are the product terms the ones we agreed, is every string still there, and does an independent model flag anything. Four passes, no fluency required.
This matters because machine translation fails quietly now. Twenty years ago bad output looked broken. Today it looks polished and confident and is occasionally about something else.
The four ways to verify an AI translation
Each pass is independent, and each produces evidence you can read in English:
Blind back-translation, to catch changed meaning.
A locked glossary with a count report, to catch terminology drift.
Mechanical diffs on segments, placeholders, and numbers, to catch omission.
A second model reviewing adversarially, to catch what the first three miss.
Run them in that order. Passes two and three are cheap enough to automate and will catch most of what actually reaches production.
Why fluency is the problem, not the solution
A modern model will give you grammatical, natural-sounding text in almost any major language. That fluency removes the signal you would normally use to detect trouble. The errors that remain are the ones that do not disturb the surface:
Meaning drift. A hedged English sentence becomes a firm promise, or vice versa.
Register slip. Your careful formal address turns into casual second person, which in German, French, Japanese, and Korean is a real error with real offence attached.
Terminology drift. Your product's "workspace" is rendered three different ways across five pages.
Silent omission. A clause, a list item, or a whole sentence disappears. Long inputs make this worse.
False friends and idioms. Rendered literally, they land somewhere between confusing and comic.
None of these look wrong to someone who cannot read the language. All of them are catchable without reading it.
Pass one: back-translation, done properly
Back-translation means translating the output back into your source language and comparing. It is the standard tool in regulated translation work, used widely in cross-cultural research and in healthcare, legal, and finance content, precisely because it surfaces meaning-level errors a fluency check misses.
The critical detail: the back-translation must be done blind. If the same conversation holds your English original in context, the model will reconstruct your English rather than translate what is actually on the page. You get a clean bill of health and learn nothing.
So use a fresh session, paste only the translated text, and ask for a literal rendering:
Translate the following text into English. Translate literally,
preserving sentence boundaries and word choice as closely as
English allows. Do not improve the phrasing. Do not add or
remove anything.
<paste target-language text only>Then compare against your original. You are looking for changed meaning, not changed wording. Back-translations always read a bit stiff. That is expected and is not itself a defect.
Be aware of the known limitation: back-translation has real shortcomings as a standalone quality test, because a good back-translator can paper over an awkward but comprehensible target text. It catches meaning errors well. It does not certify style.
Pass two: lock the glossary before you translate
Terminology drift is the most common complaint from real users of localised software, and it is entirely preventable. Decide the target-language term for every product noun before any page gets translated, then pass that list with every request.
English term | Decision | Why it is fixed |
|---|---|---|
Workspace | Keep as-is, do not translate | Matches the UI label |
Credits | Translate, single agreed term | Appears in billing copy |
Deploy | Translate, single agreed term | Users search for it |
Swarmz | Never translate | Product name |
Put it in the prompt as a hard constraint rather than a suggestion, and ask for a compliance report you can read without knowing the language:
Use this glossary exactly. After the translation, output a table
listing each glossary term, the rendering you used, and the count
of occurrences. Flag any term you could not use as specified.A count table is language-independent. If "workspace" appears eleven times in the English and the table says four, you have found a problem without reading a word of the output.
Pass three: mechanical completeness checks
These take a minute and catch the failure that embarrasses you most, which is missing content.
Segment count. Split source and target on sentence-ending punctuation and compare counts. A gap of more than about 10 percent is worth investigating. Some languages legitimately merge or split sentences, so this is a signal, not a verdict.
Placeholder integrity. If your strings contain
{{name}},%s, or<0>markers, every one of them must survive, spelled identically. Grep for them. Models rewrite placeholders more often than you would expect, and a mangled placeholder is a crash, not a typo.Number and date audit. Extract every digit sequence from both versions and compare the sets. Prices, dates, and percentages should be identical, though separators will legitimately differ, and 1,000 in English is 1.000 in German.
Length sanity. German and Finnish typically run longer than English, Chinese and Japanese much shorter. A German translation that came back shorter than the English is suspicious.
Most of this is a ten-line script. Run it in CI once you have more than a handful of strings.
Pass four: an adversarial second opinion
Use a different model, and give it a job that produces a readable-by-you answer rather than a rating.
You are reviewing a translation for a software product.
Source (English) and target are below. List every place where
the target changes the meaning, changes the level of formality,
omits content, or uses inconsistent terminology. For each, quote
the target phrase, give a literal English gloss, and say what is
wrong in one sentence. If nothing is wrong, say so.The output is in English, it is specific, and it is falsifiable. Compare it against your back-translation findings. Anything both passes flag is real. Anything only one flags is worth a human look. This is a more useful pattern than asking a model to check its own work in the same session, where it tends to agree with itself.
When to not translate at all
Some strings should stay in English, and deciding that up front saves a review cycle:
Product and feature names. Users search for them and read them in your docs. Translating "Workspace" into six languages fragments your support burden.
Error codes and technical identifiers.
ERR_RATE_LIMITis a searchable string. Localise the human-readable message beside it, not the code.Legal boilerplate you did not write. If your terms were drafted by a lawyer in one jurisdiction, a machine translation is not a legally equivalent document. Ship the English with a note, or pay for a proper translation.
Anything with a live deadline. Time-sensitive announcements are where an unreviewed error costs the most and where you have the least slack to fix it.
What none of this replaces
A native speaker, for anything that carries legal, medical, or financial weight, and for your top-of-funnel marketing copy where tone is the product. The loop above is a quality floor for interface strings, documentation, support articles, and changelogs. It is not a substitute for review when a mistranslation costs real money.
The economics are worth stating plainly. Back-translation and review cost real time, which is why translation teams reserve them for high-risk content rather than every string. Apply the full four passes to your pricing page, your onboarding, and your legal notices. Apply passes two and three to everything else.
FAQ
Can I just ask the model if its translation is accurate?
You can, and it will usually say yes. Self-assessment in the same context window is close to worthless because the model is grading a decision it just made with the original still in view. Blind back-translation and a second model are the cheap ways around this.
How long should each chunk be?
Short enough that omission is visible. Paragraph-level or string-level requests fail far less often than pasting an entire page, and they make the segment-count check meaningful. Chunking also keeps you clear of the long-input degradation that affects every model.
Do I need to keep tone consistent as well as meaning?
Yes, and it is a separate problem worth handling deliberately. The glossary handles nouns. For voice and register, give explicit instructions and examples, which is the approach in keeping tone intact when you translate.
What about right-to-left languages and scripts I cannot even segment?
The mechanical checks still work, because placeholders, digits, and glossary counts are script-independent. Segment counting needs a Unicode-aware splitter rather than a regex on full stops. Everything else in this loop is unchanged.
Is this enough to launch in a new market?
For the product surface, usually yes. For anything a customer signs or a regulator reads, no. Before you get that far, the harder work is usually structural rather than linguistic, and it is covered in wiring multi-language support into the app itself. Getting the strings right is the second problem. Getting them out of hard-coded markup is the first, and none of the verification above helps if the text is baked into a component. The general habits in the fundamentals of prompt engineering apply throughout.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


