What Are Logprobs in AI? A Practical Guide
Logprobs tell you how confident an AI model was, token by token. The arithmetic worked through, the three real uses, and the failure they never catch.
Logprobs in AI are the log of the probability a model assigned to each token it considered, returned alongside the text it generated. Where the output tells you what the model said, the logprobs tell you how confident it was in saying it, token by token. Many APIs will return them if you ask, and almost nobody does.
That is a shame, because a confidence number is the difference between an automated pipeline that knows when to stop and one that fails silently.
What logprobs in AI are, arithmetically
A model predicts the next token by producing a probability for every token in its vocabulary. Those probabilities sum to 1. The logprob is the natural logarithm of one of them.
Those raw pre-softmax scores are logits, and the probabilities come from running softmax over them. Logs are used instead of raw probabilities for a practical reason: probabilities for a long sequence get multiplied together, and multiplying many numbers below 1 underflows to zero fast. Logs turn multiplication into addition, which stays numerically stable.
The mapping is worth memorising roughly:
Probability | Logprob | Reading |
|---|---|---|
1.0 | 0.0 | Certain |
0.9 | -0.105 | Very confident |
0.5 | -0.693 | Coin flip |
0.1 | -2.303 | Unlikely but chosen |
0.01 | -4.605 | Scraping |
Logprobs are always zero or negative, and closer to zero means more confident. To get back to a probability, exponentiate: a logprob of -0.105 is exp(-0.105), which is about 0.90.
For a whole answer, sum the logprobs of its tokens, then divide by the token count to get a per-token average. Summing without dividing penalises long answers for being long, which is almost never what you want.
What it looks like in practice
Ask a model to classify a support ticket. It answers Billing. Two different worlds produce that same word:
Case A: "Billing" logprob -0.02 (probability 0.98)
next best "Account" logprob -4.1 (0.017)
Case B: "Billing" logprob -0.78 (probability 0.46)
next best "Account" logprob -0.85 (0.43)In case A the model is sure. In case B it is picking between two near-identical options and the word Billing is close to a coin flip. The text output is identical. Only the logprobs distinguish them, and only case B should be sent to a human.
This is why the gap between the top choice and the second choice is often more informative than the top probability alone. A model can be at 0.46 on the right answer and still be clearly right, if everything else is at 0.01. It can be at 0.46 and genuinely confused, if the runner-up is at 0.43.
Three real uses
Confidence gating. Set a threshold, route anything below it to review. This is the main one, and it turns a classifier from something that always answers into something that knows when to escalate. Picking the number is the hard part, covered in setting a confidence threshold for an AI classifier.
Classification without parsing. If you constrain the model to a small answer set, you can read the probabilities directly rather than parsing prose. This is the principle that purpose-built decision models make the whole product: typed answers with probabilities, no text to interpret.
Detecting guessing. Run a batch and look at the distribution of average logprobs. A cluster of low-confidence answers usually means a category your prompt does not cover, or inputs unlike anything the model has seen. That is diagnostic information you otherwise only get from users complaining.
What logprobs do not tell you
A logprob is the model's confidence in its own next token. It is not a probability that the answer is true.
A model can be extremely confident and extremely wrong. Hallucinations frequently come back with high logprobs, because the model is fluently producing text that follows from its context, and fluency is exactly what the probability measures. If the context is wrong, confident wrong text is the expected output.
So logprobs catch one failure mode well and another not at all. They catch ambiguity: the model is torn between options. They do not catch confabulation: the model is sure, and sure about something false.
A second caveat is calibration. Whether a 0.9 means the answer is right 90% of the time is a property of the specific model and your specific task, and it has to be measured. Most models are overconfident by default. The workable approach is to treat the number as an ordering rather than a probability: higher is more likely right than lower, which is enough to set a cut-off, even when the absolute value is not trustworthy.
How this relates to temperature
Temperature changes how the probabilities are sampled, not what the model believes. At temperature 0 the model takes the highest-probability token every time; higher temperatures let it pick lower-probability ones. The underlying distribution is the same.
A practical consequence: if you are using logprobs to gate decisions, run at low temperature. Otherwise you are measuring confidence in a token that was partly chosen by chance, which muddies the signal you are trying to read.
Since logprobs are per token, everything here depends on how text is split into tokens in the first place. What a tokenizer does is worth understanding if the numbers ever look strange, because a word split across three tokens has three logprobs, not one.
FAQ
How do I get logprobs from an API?
Most major APIs expose a parameter that returns them, often with an option for how many alternatives per position you want back. Check the current reference for the provider you use, since the parameter names differ and have changed over time.
Do logprobs cost extra?
Not usually in tokens. They add to the response payload, so there is bandwidth and parsing overhead, but you are not charged for information about tokens you already generated.
Can I use logprobs to detect hallucinations?
Only partially. They flag cases where the model was torn, which correlates weakly with errors. Confident hallucinations look identical to confident correct answers, so logprobs are a useful signal and not a detector.
What is a good threshold?
There is no universal number, because calibration varies by model and task. Measure it: run a few hundred labelled examples, plot accuracy against average logprob, and pick the point where accuracy drops below what you can tolerate.
Should I use the average logprob or the minimum?
The average for overall confidence, the minimum when a single uncertain token would break things, such as a number or a name in an extraction task. For extraction, one shaky token matters more than a comfortable average.
For the wider picture of what is happening underneath all this, how AI models work covers the mechanism these numbers come out of, and prompting a model to flag what it is unsure about is the prompt-level counterpart to reading the probabilities directly.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


