Dashboard

How to Prompt AI to Grade Work Against a Rubric

Ask for a score out of 10 and everything comes back a 7 or an 8. The fix is anchoring each band with a concrete example and demanding evidence before the number.

Steve Jefferson
Steve Jefferson
Developer Advocate
5 September 20261 min read

Ask a model to score something out of 10 and you get a number that moves when nothing else does. Run the same essay past it three times and you might see 7, 8, and 7. Run a slightly worse essay and you might also see 7. The scale is not measuring anything stable, because "8 out of 10" has no shared definition between you and the model.

Grading against a rubric fixes this, and the fix is not "attach a rubric." It is anchoring every band with a concrete example, forcing evidence before the score, and grading one dimension at a time.

Why a bare score is noise

Two mechanisms are at work.

The model has no reference point for the top of your scale. It has seen millions of documents rated on millions of implicit scales, and "10" is an average of all of them rather than your definition. So it regresses to the middle, which is why almost everything gets a 7 or an 8, and why the range you actually use is about three points wide.

And the score comes first. If the model outputs a number and then explains it, the explanation is a justification written after the fact, not the reasoning that produced it. You get confident prose defending an arbitrary number. Related to why the same prompt gives different answers run to run, and worse here, because a score looks precise in a way that prose does not.

The three fixes, in order of impact

1. Anchor every band with an example

This is most of the improvement. Do not write "3 = good clarity." Write what a 3 looks like, in the actual domain, in one sentence.

Compare a support-reply rubric:

text
Weak rubric:
  Clarity: 1-5, where 5 is very clear.

Anchored rubric:
  Clarity, score 1 to 4:
  1 = The customer cannot tell what will happen next.
      Example: "We are looking into this."
  2 = States an action but not who does it or when.
      Example: "This will be escalated to the right team."
  3 = States the action, the owner, and a timeframe.
      Example: "I have sent this to our billing team, who will
      reply by Thursday."
  4 = All of the above, plus what the customer should do if the
      timeframe passes.

The anchored version gives the model a target rather than a vibe. It also gives you something: when you disagree with a grade, you can point at the band and fix the wording, which you cannot do with "5 is very clear."

Note the even-numbered scale. A 1 to 4 scale has no middle to hide in. Odd scales collect a pile of 3s.

2. Evidence before score

Force the model to quote the text that justifies the grade before producing the grade.

text
For each dimension, in this order:
  1. Quote the span from the submission that most determines
     this score. Quote verbatim. If nothing in the submission
     addresses this dimension, write "no evidence".
  2. Name the rubric band this evidence matches, by its
     number and its wording.
  3. Only then, output the score.

Do not output a score for any dimension marked "no evidence".
Output "unscoreable" instead.

The "no evidence" escape hatch matters more than it looks. Without it, the model invents a score for a dimension the submission never addressed, which is the single most misleading output a grader can produce. Giving it a legitimate way to decline is what stops it from guessing.

3. One dimension per call

Grading five dimensions in one response produces halo effects: a submission that scores well on the first dimension scores well on the rest. Grading them in separate calls, each with only its own rubric band in context, removes the contamination.

This costs five times the tokens and it is worth it whenever the grade has consequences. For a rough internal triage, one call is fine.

What the full prompt looks like

Putting it together, for a single dimension:

text
You are grading one dimension of a submission against a rubric.
Grade only this dimension. Ignore spelling, length, and tone
unless the rubric mentions them.

RUBRIC
Clarity, score 1 to 4:
  [the four anchored bands from above]

SUBMISSION
[the text]

OUTPUT, in this order:
  evidence: the verbatim span that most determines the score,
            or "no evidence"
  band:     the rubric band number and its wording
  score:    the number, or "unscoreable"
  note:     one sentence, only if the submission sits between
            two bands

The note field is a small addition that pays for itself. Submissions that genuinely sit between bands are the ones where your rubric needs work, and this surfaces them instead of burying them in a rounded number. Requesting consistent output every time is easier when the shape is this rigid.

Calibrating before you trust it

Before you use the grader on anything real, grade six submissions you have already graded yourself. Pick two you consider clearly strong, two clearly weak, and two you found genuinely hard.

You are checking three things:

  • Do the clear cases match? If the model disagrees with you on an obvious strong or weak case, the rubric wording is wrong, not the model.

  • Does it separate? If all six land within one point of each other, your bands are not distinct enough. Rewrite the anchors to be further apart.

  • Is it stable? Run the same submission three times. Variation of more than one band means the rubric is still too loose.

This is a small evaluation set, and building it is the same discipline as what an AI eval is at a larger scale. When you later change the prompt or the model, rerunning these six tells you whether you improved anything, which is the point of testing whether a prompt change actually helped.

Where this genuinely does not work

Grading against a rubric works when the rubric can be written down. It fails when the quality you care about is the thing you cannot articulate: originality, taste, whether a piece of writing is actually interesting. If you find yourself writing a band that says "4 = genuinely insightful," stop. You have not defined anything, and the model will grade it as a vibe again.

The honest scope: use it for compliance-shaped judgements, where a rubric is a checklist with wording. Use a human for the rest. And never let a model be the only grader on a decision that affects a person's job or education, which is a fairness question before it is a technical one.

If you are building the rubric itself rather than applying one, prompting AI to write a hiring rubric covers the other half of the job, and the wider set of techniques lives in our guide to prompt engineering.

FAQ

What scale should I use?

Three or four bands, even-numbered. Ten-point scales give the illusion of precision that the underlying judgement cannot support.

Should I show the model examples of graded submissions?

Yes, if you have them. Two or three graded examples inside the prompt improve agreement noticeably, more than adding more rubric wording does.

Can I use a cheaper model for grading?

Often, yes. Grading against a well-anchored rubric is a much easier task than generating good work, so a smaller model frequently matches a larger one here. Test it on your calibration set before assuming either way.

How do I stop the model being too generous?

Anchor the top band with an example that is genuinely excellent rather than merely competent, and add an instruction that the top band should be rare. Generosity is usually an anchoring problem, not an instruction problem.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.