How to Use AI to Build a Decision Matrix
A full worked example of prompting AI to build a weighted decision matrix, plus the specific way models rig their own weights and how to catch it.
How to Use AI to Build a Decision Matrix
A decision matrix is a table that scores competing options against a fixed set of weighted criteria, so the choice comes out of arithmetic instead of whoever argues longest in the meeting. Using AI to build a decision matrix speeds up the parts that are tedious by hand: drafting a criteria list, running the weighted math across many rows, and pressure-testing your reasoning before you commit. It also introduces a specific risk worth knowing before you rely on it: a model can quietly settle on weights that just happen to produce whatever answer it already leaned toward. This guide covers the exact prompt to use, a full worked example with three options and five weighted criteria, and the two-step check that catches a rigged score before you act on it.
What a decision matrix actually does
A decision matrix lists your options as rows and your evaluation criteria as columns. Each criterion gets a weight reflecting how much it matters relative to the others (the weights should sum to 100%). Each option gets a raw score per criterion, usually on a 1-5 scale. Multiply score by weight, sum across criteria, and the option with the highest weighted total is the mathematically preferred choice given the inputs you supplied.
The output is only as good as two things you control: which criteria you included, and how you weighted them. AI does not fix bad inputs. What it is actually useful for is making the process of building and interrogating the matrix faster and more consistent.
Where AI helps, and where it doesn't
AI is genuinely useful for three parts of this process. First, generating a starting criteria list you might not think of unprompted, especially ones that matter but are easy to forget, like switching cost or team learning curve. Second, doing the weighted arithmetic across many rows without a transcription error, which matters more than it sounds like once you have five or six criteria and three or more options. Third, acting as a second opinion that pressure-tests a weight by asking you to justify it out loud.
AI is not useful for deciding what actually matters to your business. The weights encode your priorities and constraints, not the model's. Treat every weight the model proposes as a draft you edit, not an answer you accept.
The prompt that sets up a real decision matrix
The order of operations matters more than the wording. If you ask an AI model to build the whole matrix in a single request, options and all, it tends to infer which option you are leaning toward from context and quietly shapes the weights to support that lean. The fix is splitting the work into two prompts: one that locks in criteria and weights before the model has seen your scores, and a second that scores the options against those fixed weights.
Here is a prompt for the first step:
I need to build a weighted decision matrix for a business decision.
The decision is: [describe the decision, one sentence, no opinion
about which option you favor].
Step 1 only: propose 5 criteria I should weigh this decision on.
For each criterion, give:
- A one-sentence definition of what it measures
- A proposed weight (all weights must sum to 100%)
- A one-sentence justification for that weight that would hold
true regardless of which option wins
Do not mention or reference any specific option yet. Do not tell
me which option you think is best. Just propose criteria and
weights, with independent justification for each weight.Once you have reviewed, edited, and locked the weights, run a second prompt that hands over the fixed criteria and weights along with the actual options, and asks only for scores (1-5) and a one-line reason per cell, plus the weighted math. Keeping the two steps separate is the whole point, not a formality.
A worked example: choosing a help desk vendor
Say a support team is choosing between three help desk platforms. The names below are illustrative, not real vendors or real scores; use this as a template for the shape of the output, not as a benchmark.
The five weighted criteria, locked in before scoring:
Total cost of ownership (25%)
Implementation time (15%)
Automation depth (25%)
Reporting quality (15%)
Vendor support quality (20%)
Scoring each option 1-5 against those fixed weights produces this matrix:
Criteria (weight) | Northwind Support | Fielder Desk | Aegis Helpdesk |
|---|---|---|---|
Total cost of ownership (25%) | 5 -> 1.25 | 3 -> 0.75 | 2 -> 0.50 |
Implementation time (15%) | 4 -> 0.60 | 3 -> 0.45 | 2 -> 0.30 |
Automation depth (25%) | 2 -> 0.50 | 5 -> 1.25 | 4 -> 1.00 |
Reporting quality (15%) | 3 -> 0.45 | 4 -> 0.60 | 5 -> 0.75 |
Vendor support quality (20%) | 3 -> 0.60 | 4 -> 0.80 | 5 -> 1.00 |
Weighted total (100%) | 3.40 | 3.85 | 3.55 |
Fielder Desk comes out ahead, but only narrowly, 3.85 against Aegis Helpdesk's 3.55. That margin is itself useful information. A matrix that produces a near-tie is telling you the decision is genuinely close and probably should not be settled by the third decimal place; a matrix where one option wins on every single criterion is telling you that you probably did not need a matrix in the first place.
The failure mode: weights that produce a foregone conclusion
Here is what goes wrong when you skip the two-step split. Ask a model in one breath to “build a decision matrix to help us choose between our current vendor and two alternatives, given how frustrated the team is with the current automation,” and it will often weight automation depth heavily, because that is the framing you handed it, even if cost or implementation time actually matters more to the business this quarter. The model is not being deceptive. It is pattern-matching to “the set of weights that make this read as a well-reasoned recommendation to switch,” because that is the shape of the request you gave it.
The tell is usually subtle. The weights look individually defensible, the math is correct, and the total lands on the option that was implied as the frontrunner from the first sentence of your prompt. Nobody fabricated a number. The scoring criteria themselves were quietly built to produce that answer.
How to catch it: make it justify each weight independently
Before you let a model score anything, ask it to defend each weight in isolation, disconnected from the options. Two questions catch most rigged setups:
“Would this weight change if the options changed?” A defensible weight for automation depth should not shift just because you are comparing three different vendors instead of two. If the model's justification leans on a specific option's strengths or weaknesses rather than on what the business generally needs, the weight is contaminated.
“What would have to be true for a low-cost, low-automation option to win under these weights?” If the answer is “nothing realistic,” the weights were built to rule an outcome out, not to measure the tradeoff.
A second check worth running: after you get scores back, swap the order the options appear in and ask for a fresh scoring pass. If the weighted totals move in ways that are not explained by the actual scores, something in the framing rather than the criteria is doing the work.
Reading the result without over-trusting the number
A weighted total is a structured argument, not a verdict. Use it to see where you and a colleague actually disagree (usually it is a weight, not a score), and to catch cases where your gut preference and the math genuinely diverge, which is worth a conversation before you commit either way. Run a quick sensitivity check before finalizing anything: nudge one weight by five to ten points and see whether the ranking flips. If a small, defensible change in a single weight changes the winner, the decision is closer than the matrix makes it look, and that is worth knowing before you act on it.
This kind of two-step, weight-then-score prompting is really a specific application of a broader discipline; if you are new to structuring prompts so a model does the right work in the right order, the Prompt Engineering pillar covers the underlying framework. Teams that run this kind of analysis repeatedly also tend to save the two prompts as a reusable system prompt rather than retyping the setup each time. And if a decision made this way still goes sideways, prompting AI to write a postmortem afterward is a reasonable next step for figuring out whether the matrix or the execution was the problem.
Frequently asked questions
How many criteria should a decision matrix have?
Most business decisions are well served by four to six criteria. Fewer than that and you are probably not capturing real tradeoffs; more than that and weights get so small per criterion that the matrix stops discriminating between options in any meaningful way.
Can AI decide the weights for me?
It can propose a starting set, but the weights should reflect your organization's actual constraints and priorities, which the model does not have visibility into. Treat any AI-proposed weight as a draft to edit, and always ask for the justification behind it before accepting it.
What's the difference between a decision matrix and a pros and cons list?
A pros and cons list is unweighted and often unstructured, so a long list of minor pros can visually outweigh one major con. A decision matrix forces every factor to be scored and weighted the same way across every option, which is what makes the totals comparable.
Should every criterion use the same 1-5 scale?
Using the same scale across criteria makes the weighted math easier to sanity-check, but the scale itself matters less than consistency. What matters more is that each score has a written justification, so you can catch a criterion that got scored on vibes rather than evidence.
What if two options end up with nearly the same weighted total?
Treat it as the correct answer, not a failure of the exercise. A near-tie means the decision is genuinely close given your stated priorities, and the tie-breaker should be an explicit conversation about which risk you would rather carry, not a rerun of the matrix until something separates.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


