What Are Logits in AI? A Worked Example
Every token an AI picks starts as a raw score called a logit. Working through the arithmetic once explains temperature, top-p and log probabilities in one go.
A logit is the raw score a model assigns to a candidate next token before anything turns that score into a probability. The model produces one logit for every token in its vocabulary, all at once, and they are unbounded real numbers: 8.2, 1.4, minus 3.7. On their own they mean nothing absolute. Only the gaps between them matter. Every sampling setting you have ever adjusted, temperature included, is arithmetic performed on this list of numbers, and doing that arithmetic once by hand explains all of them.
From logits to probabilities
Suppose the model is completing "The capital of France is" and, out of a vocabulary of a hundred thousand or so, four tokens have meaningful scores:
Token | Logit |
|---|---|
Paris | 8.0 |
the | 5.0 |
located | 4.0 |
Lyon | 2.0 |
Softmax converts these into probabilities. Exponentiate each one, then divide by the total:
e^8.0 = 2981.0
e^5.0 = 148.4
e^4.0 = 54.6
e^2.0 = 7.4
--------
total = 3191.4
Paris = 2981.0 / 3191.4 = 93.4%
the = 148.4 / 3191.4 = 4.7%
located = 54.6 / 3191.4 = 1.7%
Lyon = 7.4 / 3191.4 = 0.2%Two things fall out of this that are worth carrying around. First, exponentiation is brutal: a logit gap of 3.0 becomes a twentyfold probability gap. Small differences in raw score become large differences in behaviour. Second, only relative values matter. Add 100 to every logit and the probabilities are identical, which is why an absolute logit value is not a confidence score and should never be read as one.
What temperature actually does
Temperature divides every logit before the softmax. That is the entire operation. Same four tokens, three temperatures:
Token | T = 0.5 | T = 1.0 | T = 2.0 |
|---|---|---|---|
Paris | 99.7% | 93.4% | 71.0% |
the | 0.2% | 4.7% | 15.8% |
located | 0.0% | 1.7% | 9.6% |
Lyon | 0.0% | 0.2% | 3.5% |
Dividing by 0.5 doubles every gap, so the leader runs away with it. Dividing by 2.0 halves every gap and the field bunches up. Nothing is being added or removed. The ranking never changes, only how sharply the top of the ranking dominates.
This is why temperature 0 is described as deterministic: the gaps become infinite and the highest logit always wins. It is also why high temperature produces incoherent rather than creative output past a point. At T=2.0 a token the model scored at 2.0 against a leader at 8.0 still gets picked about one time in thirty. The full behavioural picture is in what temperature in AI means.
Where top-p fits
Temperature reshapes the distribution. Top-p then truncates it. With top-p at 0.9 on the T=1.0 column above, the sampler takes tokens in descending order until their cumulative probability passes 90%: Paris alone is 93.4%, so the candidate set is one token and the other three are discarded entirely regardless of temperature.
That is the useful mental model: temperature changes the shape, top-p cuts the tail off whatever shape you produced. They interact, which is why turning both up at once produces stranger output than either alone. What top-p is goes through the sampling side in detail.
Log probabilities, and why APIs return them
Many APIs expose logprobs, the natural logarithm of each probability. Paris at 93.4% is a logprob of about minus 0.068. Less likely tokens go sharply negative: 0.2% is about minus 6.2.
Logs are used because probabilities of a long sequence multiply into numbers too small to represent, while logs add. They are also the closest thing you get to a usable confidence signal. Three practical uses:
Classification confidence. When you ask a model to output one of five labels, the logprob of the chosen label tells you whether it was a close call. A margin of 0.1 between the top two labels is a coin flip dressed as an answer.
Extraction reliability. Low logprobs on the tokens of an extracted value often mark the field the model was least sure about, which is where a human should look first.
Cheap hallucination signal. Not reliable on its own, since confidently wrong is a real state, but a sequence with unusually low average logprob is worth a second look.
Note the caveat in that last bullet. Logprobs measure the model's internal certainty about the next token, not whether the claim is true. High confidence and complete fabrication coexist comfortably, which is why checking whether an answer is hallucinated is still a separate job.
Why this matters for consistency
If you have ever wondered why the same prompt gives different answers, the logit view explains it exactly. Unless temperature is 0, you are sampling from that distribution, and any token with non-zero probability will eventually come up. A prompt where the top two logits are close is a prompt that will flip on you, and no amount of instruction fixes a near-tie. Rewriting the prompt so the correct answer wins by a wide margin is the actual fix, which is the substance of how to get consistent AI output every time.
For the layer below this, what a token in AI is covers what the vocabulary is made of, and how AI models work covers where the logits come from in the first place.
Frequently asked questions
Is a logit the same as a probability?
No. A logit is an unbounded raw score. A probability is what you get after softmax, and it is bounded between 0 and 1 and sums to 1 across the vocabulary.
Can I see logits directly?
Rarely. Most hosted APIs expose log probabilities for the top few tokens rather than raw logits. Open-weight models running locally give you the full vector.
Does a high logit mean the model is right?
It means the model finds that token likely given everything before it. That correlates with correctness on factual questions the model knows well, and not at all on questions it does not.
Why is it called a logit?
The term comes from statistics, where the logit function is the log of the odds. In neural networks it has drifted to mean the pre-softmax output of the final layer, which is related but not identical usage.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


