What Is Self-Consistency Prompting?
Self-consistency prompting samples a model multiple times and takes the majority answer. How it works, when it helps, and what it costs.
What Is Self-Consistency Prompting?
Self-consistency prompting is a technique where you ask a model to solve the same problem multiple times, using slightly different reasoning paths, then take the answer that comes up most often as the final result. Instead of trusting one chain of thought, you sample several and let majority vote filter out the reasoning paths that went wrong.
The mechanism, step by step
Ask the model to solve a problem using chain-of-thought reasoning, but generate several independent responses instead of one (most APIs let you set a sample count, or you simply repeat the same prompt N times with a non-zero temperature so the outputs vary).
Extract the final answer from each response, discarding the reasoning text.
Take the answer that the largest number of responses agreed on.
That is the entire technique. There is no additional model, no external verifier, and no extra training. The only cost is running the same prompt several times instead of once.
Why majority vote works here
Different reasoning paths through the same problem tend to fail in different, uncorrelated ways. One attempt might mis-copy a number partway through a calculation. Another might round too early. A third might apply a formula correctly and land on the right answer. If wrong answers scatter across different mistakes while correct answers converge on the same value, the value with the most votes is disproportionately likely to be the right one, even though no single attempt was verified against ground truth.
This is closest in spirit to chain-of-thought prompting, which self-consistency builds directly on top of: chain-of-thought asks for one reasoning path written out step by step, self-consistency asks for several and combines them.
A worked example
Ask a model: "A store has 3 shelves. Each shelf holds 8 boxes. Each box holds 12 items. How many items total?"
Run the same prompt five times at a moderate temperature:
Attempt | Reasoning | Answer |
|---|---|---|
1 | 3 × 8 = 24 shelves-boxes, 24 × 12 = 288 | 288 |
2 | 8 × 12 = 96 per shelf, 96 × 3 = 288 | 288 |
3 | 3 × 12 = 36, 36 × 8 = 288 | 288 |
4 | 3 × 8 = 24, 24 × 12 = 248 (arithmetic slip) | 248 |
5 | 8 × 3 = 11 (arithmetic slip), 11 × 12 = 132 | 132 |
Three different, independently-derived paths land on 288. The two wrong answers disagree with each other as much as they disagree with the right one. Majority vote picks 288, correctly, without ever checking the math against an answer key.
Where it genuinely helps, and where it does not
Self-consistency earns its extra cost on problems with a single verifiable correct answer and multiple valid paths to reach it: arithmetic, logic puzzles, multi-step reasoning with a definite answer, code that must produce a specific output. The technique needs "correct" to mean something a vote can converge on.
It does close to nothing for open-ended generation. Ask five variations of "write a tagline for a coffee shop" and you get five different, all individually reasonable, taglines with no majority to find, because there was never a single correct answer to converge toward. Applying self-consistency here just burns tokens without buying anything.
The honest cost
Running a prompt five times costs roughly five times the tokens of running it once, before any savings from caching a shared prefix. This is the same trade-off multi-agent architectures make in a different shape: spend more compute upfront in exchange for a result you trust more, and reserve the spend for problems where getting it wrong is expensive enough to justify paying for several independent tries.
A cheaper middle ground for lower-stakes problems: sample three times instead of five, or reserve the technique for the subset of a task that is actually verifiable (the final calculation) rather than applying it to an entire long, mixed reasoning-and-writing response.
Frequently asked questions
Is self-consistency the same as running a prompt twice and picking the better answer?
No. Self-consistency does not evaluate quality directly, it counts agreement across multiple independent attempts and takes whichever answer the largest number of them converged on. You never judge which single answer looks best, you count votes.
Does self-consistency require a specific model or API feature?
No special feature is required. Any model you can call multiple times with some output variation (via temperature, or simply by asking again) supports it. Some APIs offer a built-in "n" parameter to request multiple completions in one call, which is more efficient than separate requests but not required for the technique to work.
How many samples should I run?
Three to five is a common range for a meaningful majority-vote signal without excessive cost. More samples help most when answers are closely split; if three attempts already agree unanimously, running two more rarely changes the outcome.
Can self-consistency fix a model that is systematically wrong about something?
No. If a model reasons incorrectly the same way every time, for instance a consistent misunderstanding of a rule, every sample will repeat the same mistake and majority vote will confidently confirm the wrong answer. Self-consistency filters out random, uncorrelated errors. It does nothing against a consistent, repeated one.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


