How to Get AI to Flag What It's Unsure About
How to prompt AI to flag what it is unsure about: why confidence scores mislead, and three patterns that surface real uncertainty you can act on.
To get AI to flag what it is unsure about, stop asking it how sure it is. Asking for a confidence score produces a number that reads like a measurement and behaves like a guess. What works instead is structural: make the model separate what it can source from what it inferred, make it state what would change its answer, and give it an explicit route to say the answer is not available. Uncertainty becomes visible when the output format has somewhere to put it, not when you request sincerity.
Why confidence scores hide what AI is unsure about
Ask a model to rate its confidence from 0 to 100 and it will comply, fluently, forever. The research on this is unflattering. A study of verbalized confidence scores found that such scores are often poorly calibrated and that when calibration does appear it is largely incidental, varying with how the question is asked rather than with how much the model actually knows.
The mechanism is worth understanding, because it tells you what to do instead. A confidence number is generated the same way as any other token: as a plausible continuation. Humans writing confidently in the training data wrote high numbers, so the model writes high numbers. It is producing the text that usually accompanies certainty, which is not the same act as measuring its own knowledge.
The practical damage is specific. A fabricated 95 percent is worse than no number at all, because it converts a reader's healthy suspicion into misplaced trust, and it does so most reliably on exactly the confident-sounding wrong answers you most needed to catch.
That same fluency-without-grounding is the engine behind AI hallucination, and the fix has the same shape in both cases: change the structure of what you ask for, rather than asking the model to try harder.
Pattern one: source or silence
Force every claim into one of three buckets, and make the buckets part of the output format:
For each claim in your answer, label it:
[SOURCE] followed by the exact quote from the provided documents
[INFERENCE] your reasoning, stating which facts it rests on
[UNKNOWN] not answerable from what I gave you
Do not produce an unlabelled claim. If a claim would be [INFERENCE]
resting on something you cannot point to, make it [UNKNOWN] instead.This works because it does not ask for introspection at all. It asks the model to do something it is genuinely good at: locating a quote or failing to. The uncertainty surfaces as a structural fact, an empty source field, rather than as a self-assessment.
The last sentence is the load-bearing one. Without it, anything unsourced gets relabelled as inference and the bucket you cared about stays empty. Related discipline for the citation-specific case is in how to stop AI from making up citations.
Pattern two: what would change your answer
Instead of how confident are you, ask this:
End with two sections:
WOULD CHANGE MY ANSWER
Specific facts that, if different, would change your conclusion.
WHAT I ASSUMED
Anything you filled in that was not stated.This produces something you can act on. A calibration number tells you a mood. A list of assumptions tells you which three things to go and verify, and it is usually short.
It also fails informatively. When a model is genuinely on solid ground, this section is boring and specific. When it has improvised, the section is either empty, which is a tell, or it quietly admits the load-bearing guess in the third bullet. Reading that third bullet is the highest-value thirty seconds in the whole exchange.
Pattern three: make not knowing a legitimate output
Models over-answer partly because the task shape implies an answer is expected. Change the shape:
If the documents do not contain the answer, reply with exactly:
NOT IN SOURCES: <what specifically is missing>
This is a complete and correct response. Do not substitute
general knowledge for a missing source.Three details make the difference. Give the refusal an exact format, so producing it is easy rather than a deviation. Say it is correct, because otherwise the model treats it as failing the task. And name the substitution you are forbidding, since falling back on general knowledge is the specific way this goes wrong.
The mirror-image technique, asking what is absent rather than what is uncertain, is often even more effective for documents and plans. That is how to prompt AI to find what is missing.
Why models over-answer in the first place
Worth a paragraph, because it explains why all three patterns work by changing structure rather than by asking nicely. A model completing text has no separate step where it consults a store of what it knows and finds the entry empty. There is no gap to notice. The continuation that fits the question simply gets produced, and a fluent wrong answer fits a question exactly as well as a right one.
So instructions that appeal to honesty are asking for a behaviour the architecture does not have. Instructions that add a slot to the output, a source field, an assumptions list, a fixed refusal string, give the model somewhere to put the absence, and filling that slot is a text task it can do. This is also why the same wording stops working when the model changes: you are relying on the format, and formats are followed with varying enthusiasm. The related failure where an agent stops short rather than over-answers is in what model laziness is.
A worked example
Question over a folder of contracts: does our agreement with a supplier allow us to assign it to an acquirer?
The weak version asks, with a confidence rating. The answer comes back: yes, assignment is permitted on notice, confidence 90 percent. Clean, quotable, and possibly invented.
The structured version, using patterns one and two, returns something like this:
[SOURCE] Clause 14.2: this Agreement may be assigned with the prior written consent of the Supplier.
[INFERENCE] Consent is therefore required, resting on clause 14.2 being the only assignment provision.
[UNKNOWN] Whether consent may be withheld unreasonably. No standard is stated in the provided text.
WOULD CHANGE MY ANSWER: a change-of-control clause elsewhere in the contract, or a side letter. I was given clauses 1 to 20 only.
The second answer is less satisfying and far more useful. It has converted a false yes into one precise question for a lawyer, and it told you the range of what it read.
How to check it is actually working
Do not take the format at face value. Two cheap tests:
Ask something you know is not in the sources. If you get a fluent answer instead of the not-in-sources form, the instruction is not binding and needs to move earlier or be repeated in the output format.
Ask something you know is in the sources and verify the quote character by character. A model that paraphrases inside a [SOURCE] label has given you the appearance of grounding, which is the worst of both worlds.
Run both whenever you change the prompt. The behaviour is not stable across model versions, so a pattern that held last quarter deserves re-testing. There is a complementary self-review step in how to prompt AI to check its own work, and the wider technique set is in our prompt engineering guide.
FAQ
Can I trust an AI model's confidence percentage?
Not as a measurement. Research on verbalized confidence finds these scores are frequently poorly calibrated and sensitive to phrasing, with models tending to overconfidence. Use them at most as a weak signal, never as a threshold for whether to check something.
How do I get AI to say I do not know?
Give not knowing an exact output format, state explicitly that it is a correct response, and forbid the specific substitution you are worried about, usually falling back on general knowledge when a source is missing. Models under-use refusals mostly because the task shape implies an answer is required.
Does asking are you sure improve accuracy?
Rarely, and it can make things worse. It often prompts a revision that sounds more certain, or a reversal of a correct answer under social pressure. Asking what would change the answer is more productive because it requests specifics instead of a verdict.
Should I use structured output instead of prompt instructions?
Where your tooling supports it, yes. A schema with a required sources array is harder to ignore than a sentence, and an empty array is unambiguous. The patterns here are the same idea expressed in prose for tools without schema support.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


