Dashboard

What Is Constitutional AI?

Constitutional AI replaces most of the human labelling in alignment training with a written document the model critiques itself against. Here is the mechanism.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
18 September 20261 min read

What Is Constitutional AI?

Constitutional AI is a training method where a model is shaped by a written set of principles, a constitution, rather than by human ratings of individual responses. The model generates an answer, criticizes its own answer against the principles, rewrites it, and learns from the rewrite. Humans write the document. The model does the labelling.

The name is literal. There is an actual text, and the model is trained to behave consistently with it.

Why it exists

Reinforcement learning from human feedback works and does not scale gracefully. Every preference signal costs a human sitting down and choosing between two responses, tens of thousands of times, on material that is frequently unpleasant. It is slow, expensive, and inconsistent, because human raters disagree with each other and with themselves on Tuesdays.

Constitutional AI replaces most of that labelling with model-generated preferences guided by written rules. The original Anthropic paper, Constitutional AI: Harmlessness from AI Feedback, puts the design goal plainly: the only human oversight is a list of principles.

If RLHF is the baseline, this is the variant where the preference signal comes from a model reading rules.

The two phases

Phase one: supervised self-critique

The model is prompted, usually with something it should decline or handle carefully. It produces an initial response. Then it is asked to critique that response against a principle sampled from the constitution, and to revise it accordingly. The revision, not the original, becomes training data, and the model is fine-tuned on those revisions.

The loop is: answer, criticize, rewrite, learn from the rewrite.

Phase two: reinforcement learning from AI feedback

The fine-tuned model generates pairs of responses. A model, not a person, evaluates which one better satisfies principles sampled from the constitution. Those AI-generated preferences train a reward model, and the reward model drives reinforcement learning exactly as it would in RLHF.

That substitution is what people mean by RLAIF: reinforcement learning from AI feedback. The machinery is the same as RLHF. The source of the preference signal is different.

What is in a constitution

The document has grown considerably. Anthropic's version in 2022 was roughly a page of direct instructions along the lines of do not be racist, toxic, or violent, with the UN Declaration of Human Rights providing the broad ethical footing. The current version runs to roughly a hundred pages and reads less like a content policy and more like a document about ethics, autonomy and judgment under uncertainty.

That growth is the interesting trend. A short list of prohibitions produces a model that follows a short list of prohibitions. A longer document describing how to weigh competing considerations produces something that behaves more like it is reasoning about the situation, with all the unpredictability that implies. The current constitution's treatment of the model's own moral status is contested enough to have produced a public disagreement between Microsoft AI and Anthropic this month.

What it is good at, and what it is not

Property

Constitutional AI

Plain RLHF

Human labels needed

Few

Many

Rules are inspectable

Yes, written down

No, implicit in ratings

Consistency across cases

High

Varies by rater

Captures unstated human taste

Poorly

Well

Cost to change behaviour

Edit the document, retrain

Re-label, retrain

The inspectability is the underrated advantage. When a model trained this way refuses something, there is a document you can point at, which is a very different situation from a refusal that emerged from a preference dataset nobody can read. If you have wondered why models refuse some requests and not others, this is part of the answer for models trained this way.

The weakness is the mirror image. A written constitution captures what someone thought to write down. Human raters, for all their inconsistency, carry an enormous amount of unstated judgment about what a good response feels like, and a document does not.

Is this only an Anthropic thing

The specific constitution is Anthropic's. The technique is not proprietary, and AI-generated preference data has become common across the industry as part of post-training generally. Most frontier labs now use some mixture of human and model-generated preference signal, whether or not they call the rules a constitution.

What remains relatively distinctive is publishing the document. A model card tells you what a model scored. A published constitution tells you what it was aiming at, which is a different and in some ways more useful disclosure. Our explainer on what an AI model card contains covers the other half of that picture.

FAQ

What is the difference between RLHF and RLAIF?

The pipeline is nearly identical. RLHF trains a reward model on human preferences between responses. RLAIF trains it on preferences generated by a model following written principles. The humans move from rating responses to writing rules.

Does a constitution mean the model always follows the rules?

No. It shapes training, it does not enforce behaviour at runtime. A model trained against a constitution still has failure modes, and jailbreaks still exist.

Can I write a constitution for my own application?

Not in the training sense without training a model. You can do the shallow version, which is a detailed system prompt describing principles and asking the model to check its draft against them before answering. That is the same shape at a much smaller scale.

Why has the constitution grown from one page to a hundred?

Short prohibition lists produce brittle behaviour on cases they did not anticipate. A longer document describing how to weigh competing considerations generalizes better to situations nobody wrote down, at the cost of being harder to predict.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.