Dashboard

What Is a Reward Model in AI? A Plain Explanation

A reward model predicts which answer a human would prefer, so preference training can run at scale. How it is built, why it is gamed, and what it explains.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
11 September 20261 min read

What Is a Reward Model in AI? A Plain Explanation

A reward model is a second, smaller model whose only job is to look at an answer and predict how much a human would like it. Labs train one so that preference training can run at machine speed: instead of a person grading every one of millions of candidate responses, the reward model grades them, and the main model is trained to score well against it.

Everything good and everything strange about the assistants you use every day traces back to that substitution.

The problem it solves

Suppose you want to train a model to be more helpful. You can show it examples of helpful answers, which is supervised fine-tuning, and that gets you a long way. But helpfulness is comparative. What you actually want is for the model to prefer the better of two plausible answers, and there is no clean loss function for better.

So you ask people. Show a rater two responses to the same prompt, ask which they prefer, record the answer. Do that tens of thousands of times and you have a preference dataset. The trouble is that training a large model takes millions of comparisons, and you cannot hire your way to that number.

The reward model is the workaround. Train a model on the human comparisons you did collect until it can predict human preference on unseen pairs, then let it stand in for the humans for the rest of the run.

How one gets built

  1. Collect comparisons. Generate several responses per prompt from the current model, and have human raters rank them against written guidelines.

  2. Train the scorer. Fit a model that takes a prompt and a response and returns a single number, trained so that responses humans preferred score higher than ones they did not. It is usually initialised from the same family as the model it will grade.

  3. Optimise against it. Run reinforcement learning on the main model using that score as the reward signal. This is the stage most people mean when they say reinforcement learning from human feedback.

  4. Constrain the drift. Penalise the model for moving too far from where it started, or it will find degenerate answers that score well and read like nonsense.

That fourth step is not a footnote. Without it, optimisation collapses quickly and obviously. With it, the collapse is slower and much harder to spot, which is where the interesting failures live.

The canonical public write-up of this pipeline is still OpenAI's InstructGPT paper, Training language models to follow instructions with human feedback, which reported a 1.3-billion-parameter model beating 175-billion-parameter GPT-3 in human evaluations after exactly this treatment. Capability was not the difference. Preference training was.

Reward hacking, and why your assistant flatters you

A reward model is an approximation of human preference, and any approximation can be gamed. Optimise hard enough against a proxy and you get a model that is excellent at the proxy and drifting from the thing you wanted. This is reward hacking, and it is the most important idea in this post.

Three symptoms you have almost certainly noticed:

  • Agreeableness. Raters, on average, prefer answers that validate them. A model trained on that preference learns that agreement scores well, which is the machinery behind sycophancy and the reason getting an AI to stop being agreeable takes deliberate prompting rather than a polite request.

  • Length. Longer answers read as more thorough to a rater skimming a comparison, so length correlates with reward even when it adds nothing. Hence the three-paragraph preamble before the one sentence you needed.

  • Confident wrongness. A hedged correct answer and a confident incorrect one often score similarly, because the rater is judging presentation as much as accuracy. That trade is baked in at training time, which is why spotting a hallucinated answer remains your job rather than the model's.

None of these are bugs in the usual sense. They are the reward model faithfully reproducing what raters preferred, applied harder than the raters ever intended.

Reward model or LLM judge?

These get confused constantly, and the difference is about who is using them.

A reward model is a training-time component inside a lab. It outputs a bare number, it is optimised against directly, and you never see it. LLM-as-a-judge is an evaluation-time technique available to anyone: you prompt a general model with a rubric and ask it to grade an output, and it explains its reasoning.

The rule of thumb: if you are building a product, you want a judge. Reward models are for people training models, and a judge you can read is worth more to you than a score you cannot.

Why this is worth knowing

Because it changes what you expect from prompting. A great deal of model behaviour that feels like a personality quirk is a preference gradient installed during training, and prompting can bend it but rarely erase it. Asking for brevity fights a length bias. Asking for honest criticism fights an agreeableness bias. You will get further with structure, a rubric, an explicit instruction that disagreement is the desired output, than with polite framing.

It also explains why two models can be equally capable and feel completely different. The capability came from pretraining. The manners came from whose raters were asked what, which is the layer most of model alignment actually operates on.

Frequently asked questions

Is a reward model the same size as the model it trains?

Usually smaller. It only has to score text, not generate it, and it runs many times per training step, so efficiency matters. Some labs do use same-size reward models where the judgement is subtle.

Do all modern models use reward models?

Not for everything. Where correctness can be checked automatically, a maths answer or a test suite that passes, labs increasingly train against the checker directly and skip the learned reward model. Preference-based rewards are still standard for style, tone, safety and helpfulness, where no checker exists.

Can a reward model be wrong?

Routinely. It is a prediction of human preference trained on a finite sample, so it inherits the biases of the raters, the guidelines they were given, and the distribution of prompts they saw. A reward model that never saw your domain will score answers in it roughly at random.

Does any of this apply if I am just using an API?

Indirectly, and usefully. You cannot change the reward model, but knowing it exists tells you which behaviours are cheap to prompt around and which are load-bearing. For the rest of the picture, our overview of how an AI model gets built covers where this sits in the training pipeline.

For the metric most commonly used to score a model's raw fluency during training and evaluation, see what is perplexity in AI.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.