Dashboard

How to Prompt AI for a Confidence Score Worth Using

Rate your confidence out of 10 gives you a meaningless number. Evidence-anchored bands, a worked contract example, and a 20-minute calibration test.

Steve Jefferson
Steve Jefferson
Developer Advocate
11 September 20261 min read

How to Prompt AI for a Confidence Score Worth Using

Asking a model to rate its confidence from 1 to 10 gives you a number, and that number is close to useless. You will get an 8 for almost everything, a 9 when it is wrong in a fluent way, and a 6 when your question was phrased hesitantly. To get a confidence score you can actually route on, you have to stop asking for a feeling and start asking for a verdict against defined evidence bands, then spend twenty minutes checking whether the bands hold up on your own data.

Below is the prompt pattern, a worked example, and the calibration test that tells you whether to trust the output.

Why the obvious version fails

Three separate problems, and they compound.

There is no anchor. A 7 out of 10 means nothing without a definition of what a 4 or a 9 would look like. The model has to invent a scale for every request, and it invents a different one each time. Two answers scored 7 in the same session are not comparable to each other, let alone across sessions.

The score tracks tone, not evidence. Confidence in the output correlates with how confidently the answer was phrased, which is itself a product of training preferences rather than any check against the source. You are measuring fluency wearing a numeric hat.

Being wrong is free. Nothing in the prompt makes a high score costly. A model that says 9 and is wrong pays nothing, so there is no pressure toward honesty. This is the same mechanism that makes getting a model to say it does not know harder than it should be.

The fix: bands tied to evidence, not feeling

Replace the open scale with a small number of named levels, each defined by what must be true in the source material. Four levels is the sweet spot. Ten is too many to distinguish, two throws away information.

For each answer, assign exactly one confidence level and state it before the answer:

CERTAIN - the answer is stated explicitly in the provided source. Quote the sentence.

INFERRED - the answer follows from two or more statements in the source, but is not stated directly. Name the statements.

GUESSED - the answer draws on general knowledge outside the provided source. Say what you are assuming.

UNAVAILABLE - the source does not contain enough to answer. Do not guess.

You may not use CERTAIN unless you can quote the exact supporting text.

The last line is the load-bearing one. It attaches a cost to the top band: claiming it obliges the model to produce a quotation, and a fabricated quotation is something you can catch automatically. The score stops being an opinion and becomes a claim you can verify.

Worked example: pulling a payment term out of a contract

Task: extract the payment window from a supplier agreement. Here is the difference in practice.

Vague prompt: What are the payment terms? Rate your confidence out of 10.

Payment terms are net 30 days from invoice date. Confidence: 8/10.

The 8 is doing no work. It is equally consistent with the clause being on page 4 and with the model pattern-matching the most common payment term in the world.

Banded prompt, same document:

CERTAIN. Quote: "Client shall pay all undisputed invoices within thirty (30) days of receipt." Answer: net 30 from receipt of invoice, not from invoice date.

Two things improved. You can verify the claim in four seconds by searching the document for that string. And the band forced a distinction the vague answer flattened: from receipt, not from invoice date. That difference is worth real money on a late-paying account, and the first answer got it wrong while sounding confident.

When the clause genuinely is not there, a banded prompt returns UNAVAILABLE rather than a plausible invention, which is most of the value. Pair it with prompting the model to find what is missing if absence is what you care about.

Test the bands before you trust them

A confidence score is only worth having if high confidence is actually more accurate than low confidence. That is a measurable property and it takes about twenty minutes to check.

  1. Take 20 to 30 real items from your workload, not invented examples. You need ones where you already know the right answer.

  2. Run them through the banded prompt.

  3. Bucket the results by the band the model assigned.

  4. Within each bucket, count how many answers were actually correct.

What you are looking for is separation. If CERTAIN comes out at 95 percent correct and GUESSED at 60 percent, the bands are informative and you can route on them. If CERTAIN and GUESSED both land near 80 percent, the model is assigning bands at random and the score is decoration. Tighten the definitions, or accept that this task does not support self-assessment and route everything to review.

Two failure patterns worth naming. If almost everything comes back CERTAIN, your top band is too easy to claim and the quotation requirement is probably not being enforced. If almost nothing does, the model is being defensive and you may need an explicit instruction that UNAVAILABLE is a normal outcome rather than a failure. This is a small evaluation set and rerunning it is how you know a prompt change actually helped.

When to use token probabilities instead

If you have API access and the task reduces to a short answer from a known set, a classification, a field value, a yes or no, the model's own token probabilities are a better confidence signal than anything it says about itself. They come from the generation process rather than from a second act of self-description, and they are not subject to the tone bias above.

Two practical notes. OpenAI exposes this through the logprobs and top_logprobs parameters on chat completions. And support is not universal across a provider's own range: OpenAI's changelog records that GPT-6 Astra does not support custom temperature, top_p or log probabilities at all, so a pipeline built on token probabilities can break when you upgrade the model underneath it.

They have limits. They tell you the model was confident in the tokens it emitted, which is not the same as those tokens being true, and they are meaningless for long free-form answers where confidence varies sentence by sentence. For extraction and classification they are excellent. For an essay, use bands.

Doing something with the number

A score nobody acts on is theatre. Wire it to a decision before you ship it:

  • CERTAIN and INFERRED with a named source: auto-accept, log the quotation alongside the value.

  • GUESSED: hold for human review, and show the stated assumption in the review queue so the reviewer knows what to check.

  • UNAVAILABLE: route back to the source. This is a data-collection problem, not a model problem, and treating it as one stops people from prompting harder at a document that does not contain the answer.

Set the thresholds from your calibration run rather than from instinct, and rerun the check whenever you change models. Band behaviour does not transfer reliably between model families, and a prompt that was well calibrated on one can be badly calibrated on the next.

Frequently asked questions

Why not just ask for a percentage?

Percentages look precise and encourage false precision. A model producing 73 percent is not distinguishing that from 76 percent in any meaningful way. Named bands with written definitions carry more information and are easier to audit.

Does asking for reasoning first improve the score?

Usually yes, for tasks with an inference step. Having the model lay out its support before committing to a band gives the band something to be about. The pattern overlaps with prompting a model to explain its reasoning before answering, and the same caution applies: stated reasoning is a useful artefact, not a transcript of what the model did internally.

Can I trust confidence scores on subjective tasks?

Less so. Bands work because they point at evidence. On a task where there is no ground truth to point at, such as whether a piece of copy is on brand, you are back to eliciting an opinion. Use a rubric and a human spot check instead.

Should the model score its own output or should a second call do it?

A second call with only the output and the source, not the original reasoning, is more honest. Self-scoring in the same breath tends to justify what was just written. This is the same reasoning behind prompting a model to check its own work in a separate pass.

For the broader set of techniques this builds on, see our prompt engineering guide.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.