Dashboard

How to Set a Confidence Threshold for an AI Classifier

A confidence threshold decides which AI answers run automatically and which go to a person. Set it from a labelled test set, then price errors against review time.

Steve Jefferson
Steve Jefferson
Developer Advocate
30 September 20261 min read

To set a confidence threshold for an AI classifier, run it on 100 to 200 real items that a person has already labelled, record the confidence it reports for each answer, then test several cut-offs and price two costs at each one: the mistakes that slip through automatically and the items you send to a human. Pick the cut-off where the total is lowest. Nothing about 0.5, 0.8 or 0.9 is special. The right number depends on what a wrong answer costs you.

This guide walks through the whole process with a worked example. The numbers in it are invented to show the arithmetic, not measurements from any real system, and every figure adds up so you can follow it by hand.

Why a default threshold is a guess

A classifier gives you a label and usually a score. The temptation is to pick a round number, automate everything above it, and move on. That fails in two quiet ways. If the threshold is too low, wrong answers flow through unchecked and you find out from an angry customer. If it is too high, almost everything lands in a review queue and the automation saves nobody any time.

The threshold is an economic decision disguised as a technical one. It trades two kinds of cost against each other, and only your own data can tell you the exchange rate.

Step 1: build a labelled test set

You need real inputs, not invented ones. Pull 100 to 200 recent items from wherever the classifier will work, such as support messages, form submissions or uploaded documents. Have a person assign the correct label to each one, without looking at what the model said.

Three rules keep the test honest:

  • Include the awkward cases on purpose. Messages that could fit two categories, very short ones, and ones in a second language belong in the set.

  • Keep it separate from anything you used to write the prompt. If the model has seen these examples, the test flatters it.

  • Write down how you labelled ambiguous items. You will need the rule again when you re-test.

If you are still deciding how to word the classification prompt itself, start with our guide to prompting AI to triage support tickets, then come back here.

Step 2: record the score for every item

Run the classifier over the test set and save four things per item: the text, the true label, the predicted label and the confidence. A plain CSV is fine.

Where the confidence comes from matters. A score you get by asking a general model to state how sure it is tends to be weaker than a score produced by the classifier itself, which we cover in how to prompt AI for a confidence score worth using. Some purpose-built classification endpoints return a score alongside the chosen answer. OpenAI's new Decisions API is reported to work that way, as opentools.ai describes, and The Decoder reports it is in limited preview. Our Decisions API report separates what is confirmed from what is not. Whatever the source, the method below is the same.

Step 3: sweep the confidence threshold

The script below reads the CSV and, for each candidate threshold, counts how many items would be handled automatically, how many of those would be wrong, and how many would go to review.

python
import csv

# columns: text, true, predicted, confidence (0 to 1)
rows = list(csv.DictReader(open("labeled_run.csv")))

def sweep(rows, thresholds, minutes_per_error=10, minutes_per_review=1):
    for t in thresholds:
        auto = [r for r in rows if float(r["confidence"]) >= t]
        wrong = sum(r["predicted"] != r["true"] for r in auto)
        review = len(rows) - len(auto)
        minutes = wrong * minutes_per_error + review * minutes_per_review
        print(f"{t:.2f}  auto={len(auto):3d}  wrong={wrong:2d}  "
              f"review={review:3d}  minutes={minutes}")

sweep(rows, [0.50, 0.70, 0.85, 0.95])

Here is what a run might print for an invented set of 200 messages:

Threshold

Handled automatically

Wrong among those

Error rate

Sent to review

0.50

190

19

10.0%

10

0.70

160

8

5.0%

40

0.85

120

2

1.7%

80

0.95

60

0

0%

140

The pattern is general. Raising the threshold always shrinks the automatic pile and lowers its error rate, and always grows the review pile.

Step 4: price the two kinds of mistake

Now attach a cost to each. Suppose a wrong automatic answer costs 10 minutes to find and fix, and reviewing one item costs 1 minute. The last column of the sweep multiplies it out.

Threshold

Errors x 10 min

Reviews x 1 min

Total minutes

0.50

190

10

200

0.70

80

40

120

0.85

20

80

100

0.95

0

140

140

In this example 0.85 wins at 100 minutes. The lowest threshold is worst because it leaves too many errors, and the highest is worse than 0.85 because it pays for 140 reviews to prevent errors that cost only 20 minutes.

Now change one assumption. Say a wrong answer is expensive, such as a refund sent to the wrong customer, and costs 60 minutes of cleanup. The same four rows give 0.50 at 1,150 minutes, 0.70 at 520, 0.85 at 200 and 0.95 at 140. The best threshold jumps to 0.95. Same classifier, same data, different winner, because the price of a mistake changed. That is why borrowing someone else's cut-off is a bad idea.

Step 5: check that the scores mean what they say

A threshold only works if a higher score really does mean a more likely correct answer. Test it by grouping the test items into confidence bands, for example 0.5 to 0.7, 0.7 to 0.9 and 0.9 to 1.0, and computing the share correct in each band.

  • If accuracy rises steadily with each band, the score is usable.

  • If the top band is barely more accurate than the middle band, the score is noise and no threshold will help. Use a different score source or route everything to review.

  • If the top band is correct far less often than its score implies, the model is overconfident. Raise your threshold to compensate, and trust the measured accuracy, not the stated number.

Step 6: design the review queue

The queue is half the system, and a bad one erases the savings. Four practices help:

  1. Show the model's guess next to the item so the reviewer confirms or corrects in one click instead of starting from scratch.

  2. Sort by lowest confidence first, because those are the items most likely to need a real decision.

  3. Record every correction. Corrected items become new labelled data for your next test set.

  4. Sample a small share of automatically handled items each week and check them anyway. It is the only way to notice the classifier drifting while the queue looks quiet.

Step 7: re-test on a schedule

A threshold is tied to a model, a prompt and a mix of inputs. Re-run the sweep when any of those change: a new product line that brings different messages, a new category added to the list, or a swap to another model. The process for comparing models on your own data is in how to test a new AI model before switching, and the general idea of scoring outputs against known answers is covered in what an AI eval is. If you are wiring all of this into something you are building, the wider picture is in our guide to building an app with AI.

Common mistakes

  • Choosing the threshold on the same examples used to write the prompt.

  • Using one threshold for every category. A wrong "billing" answer and a wrong "spam" answer rarely cost the same, so set a separate threshold per label if the costs differ.

  • Ignoring the review pile. A threshold that sends 70 percent of items to people is an expensive search box.

  • Never re-testing after the inputs change.

FAQ

What is a good confidence threshold for an AI classifier?

There is no universal number. Measure it: sweep several thresholds on a labelled test set and pick the one with the lowest combined cost of automatic errors and human review time. A costly mistake pushes the answer higher.

How many labelled examples do I need?

100 to 200 real items is enough to see the shape of the trade-off for a first decision. More helps if some categories are rare, because a band with only a handful of items tells you very little.

What if the model's confidence scores are all high?

Check calibration by band. If accuracy barely changes between bands, the scores are not informative, and you should either change where the score comes from or send more items to review.

Should every category share one threshold?

Only if the cost of a wrong answer is the same across categories. Where it differs, run the sweep per category and set separate cut-offs.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.