How to Prompt AI to Grade Work Against a Rubric
Ask for a score out of 10 and everything comes back a 7 or an 8. The fix is anchoring each band with a concrete example and demanding evidence before the number.
Ask a model to score something out of 10 and you get a number that moves when nothing else does. Run the same essay past it three times and you might see 7, 8, and 7. Run a slightly worse essay and you might also see 7. The scale is not measuring anything stable, because "8 out of 10" has no shared definition between you and the model.
Grading against a rubric fixes this, and the fix is not "attach a rubric." It is anchoring every band with a concrete example, forcing evidence before the score, and grading one dimension at a time.
Why a bare score is noise
Two mechanisms are at work.
The model has no reference point for the top of your scale. It has seen millions of documents rated on millions of implicit scales, and "10" is an average of all of them rather than your definition. So it regresses to the middle, which is why almost everything gets a 7 or an 8, and why the range you actually use is about three points wide.
And the score comes first. If the model outputs a number and then explains it, the explanation is a justification written after the fact, not the reasoning that produced it. You get confident prose defending an arbitrary number. Related to why the same prompt gives different answers run to run, and worse here, because a score looks precise in a way that prose does not.
The three fixes, in order of impact
1. Anchor every band with an example
This is most of the improvement. Do not write "3 = good clarity." Write what a 3 looks like, in the actual domain, in one sentence.
Compare a support-reply rubric:
Weak rubric:
Clarity: 1-5, where 5 is very clear.
Anchored rubric:
Clarity, score 1 to 4:
1 = The customer cannot tell what will happen next.
Example: "We are looking into this."
2 = States an action but not who does it or when.
Example: "This will be escalated to the right team."
3 = States the action, the owner, and a timeframe.
Example: "I have sent this to our billing team, who will
reply by Thursday."
4 = All of the above, plus what the customer should do if the
timeframe passes.The anchored version gives the model a target rather than a vibe. It also gives you something: when you disagree with a grade, you can point at the band and fix the wording, which you cannot do with "5 is very clear."
Note the even-numbered scale. A 1 to 4 scale has no middle to hide in. Odd scales collect a pile of 3s.
2. Evidence before score
Force the model to quote the text that justifies the grade before producing the grade.
For each dimension, in this order:
1. Quote the span from the submission that most determines
this score. Quote verbatim. If nothing in the submission
addresses this dimension, write "no evidence".
2. Name the rubric band this evidence matches, by its
number and its wording.
3. Only then, output the score.
Do not output a score for any dimension marked "no evidence".
Output "unscoreable" instead.The "no evidence" escape hatch matters more than it looks. Without it, the model invents a score for a dimension the submission never addressed, which is the single most misleading output a grader can produce. Giving it a legitimate way to decline is what stops it from guessing.
3. One dimension per call
Grading five dimensions in one response produces halo effects: a submission that scores well on the first dimension scores well on the rest. Grading them in separate calls, each with only its own rubric band in context, removes the contamination.
This costs five times the tokens and it is worth it whenever the grade has consequences. For a rough internal triage, one call is fine.
What the full prompt looks like
Putting it together, for a single dimension:
You are grading one dimension of a submission against a rubric.
Grade only this dimension. Ignore spelling, length, and tone
unless the rubric mentions them.
RUBRIC
Clarity, score 1 to 4:
[the four anchored bands from above]
SUBMISSION
[the text]
OUTPUT, in this order:
evidence: the verbatim span that most determines the score,
or "no evidence"
band: the rubric band number and its wording
score: the number, or "unscoreable"
note: one sentence, only if the submission sits between
two bandsThe note field is a small addition that pays for itself. Submissions that genuinely sit between bands are the ones where your rubric needs work, and this surfaces them instead of burying them in a rounded number. Requesting consistent output every time is easier when the shape is this rigid.
Calibrating before you trust it
Before you use the grader on anything real, grade six submissions you have already graded yourself. Pick two you consider clearly strong, two clearly weak, and two you found genuinely hard.
You are checking three things:
Do the clear cases match? If the model disagrees with you on an obvious strong or weak case, the rubric wording is wrong, not the model.
Does it separate? If all six land within one point of each other, your bands are not distinct enough. Rewrite the anchors to be further apart.
Is it stable? Run the same submission three times. Variation of more than one band means the rubric is still too loose.
This is a small evaluation set, and building it is the same discipline as what an AI eval is at a larger scale. When you later change the prompt or the model, rerunning these six tells you whether you improved anything, which is the point of testing whether a prompt change actually helped.
Where this genuinely does not work
Grading against a rubric works when the rubric can be written down. It fails when the quality you care about is the thing you cannot articulate: originality, taste, whether a piece of writing is actually interesting. If you find yourself writing a band that says "4 = genuinely insightful," stop. You have not defined anything, and the model will grade it as a vibe again.
The honest scope: use it for compliance-shaped judgements, where a rubric is a checklist with wording. Use a human for the rest. And never let a model be the only grader on a decision that affects a person's job or education, which is a fairness question before it is a technical one.
If you are building the rubric itself rather than applying one, prompting AI to write a hiring rubric covers the other half of the job, and the wider set of techniques lives in our guide to prompt engineering.
FAQ
What scale should I use?
Three or four bands, even-numbered. Ten-point scales give the illusion of precision that the underlying judgement cannot support.
Should I show the model examples of graded submissions?
Yes, if you have them. Two or three graded examples inside the prompt improve agreement noticeably, more than adding more rubric wording does.
Can I use a cheaper model for grading?
Often, yes. Grading against a well-anchored rubric is a much easier task than generating good work, so a smaller model frequently matches a larger one here. Test it on your calibration set before assuming either way.
How do I stop the model being too generous?
Anchor the top band with an example that is genuinely excellent rather than merely competent, and add an instruction that the top band should be rare. Generosity is usually an anchoring problem, not an instruction problem.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


