Base Model vs Instruct Model: The Difference
A base model predicts what text comes next. An instruct model has been trained to treat your text as a request. Same weights underneath, very different behaviour.
Base Model vs Instruct Model: What's the Difference?
A base model predicts what text comes next. An instruct model has been trained on top of a base model to treat your text as a request and respond to it. Same weights underneath, very different behaviour, and the gap explains most of the confusion people hit the first time they download an open-weight model and it refuses to act like a chatbot.
Ask a base model "What is the capital of France?" and a plausible continuation is another question, because in the training data that sentence often appears in a list of quiz questions. Ask an instruct model the same thing and you get "Paris". Neither is broken. One is completing a document, the other has been taught that a document ending in a question is a request for an answer.
The pipeline that separates them
Every instruct model starts life as a base model. The stages between them are roughly these:
Pretraining. The model is trained to predict the next token across an enormous corpus. What comes out is a base model, sometimes called a foundation or pretrained model. It has absorbed grammar, facts, code, and reasoning patterns, and it has no notion of a conversation, a user, or a task. Our explainer on how AI models work covers this stage in more depth.
Supervised fine-tuning. The base model is trained further on curated examples of instructions paired with good responses. This is where the request-and-response shape gets installed. It is ordinary fine-tuning, applied to a specific kind of data.
Preference training. The model is then tuned against human or model judgements of which of two responses is better, which is where refusal behaviour, tone, and formatting habits mostly come from. RLHF is the best known version of this stage.
Vendors label the results differently. You will see -base, -pt or no suffix for the first, and -instruct, -it, -chat or -sft for the later stages. The suffix is the only reliable signal on a model listing, and it is worth checking before you download 40GB.
What changes in practice
Behaviour | Base model | Instruct model |
|---|---|---|
Response to a question | Often continues the document, may ask more questions | Answers the question |
Stopping | Runs until it hits your token limit | Emits an end-of-turn token and stops |
System prompts | No concept of them | Understands and generally follows them |
Refusals | Almost none | Trained refusal behaviour |
Chat formatting | None. You supply the structure | Expects a specific template |
Few-shot examples | Very responsive to them | Responsive, but often unnecessary |
The stopping behaviour is the one that surprises people most. A base model has no reason to stop, because documents do not end after one paragraph. If you run one without a stop sequence, it will happily generate a fake follow-up question, then a fake answer, then a fake email signature.
Chat templates are not optional
An instruct model is trained with specific control tokens marking who is speaking. Hugging Face's chat templating documentation gives the canonical example: Mistral-7B-Instruct uses [INST] and [/INST] around user messages, while Zephyr-7B, fine-tuned from the same base model, uses <|user|> and <|assistant|> role markers. Feed one model the other's format and quality drops noticeably, because the model is now seeing a token pattern it was never trained on.
This is why apply_chat_template() exists rather than string concatenation. The chat basics guide walks through passing a list of role and content dictionaries and letting the tokeniser insert the right markers. If you are calling a hosted API instead of running weights yourself, the provider does this for you, which is exactly why the distinction is invisible until the first time you run something locally.
When you would actually reach for a base model
Most people never should. The two real cases:
You are fine-tuning. If you plan to train on your own instruction data, starting from a base model avoids fighting an existing instruction style. Starting from an instruct model is also viable and usually cheaper in data, but you inherit its formatting habits and refusal behaviour along with everything else.
You are doing pure completion work. Autocomplete inside a code editor, template filling, or continuing a partial document are completion tasks, not conversational ones. A base model is often better at them because it has not been taught to preface everything with a friendly summary.
Everything else, including nearly all agent work, retrieval pipelines, and classification, wants an instruct model. If your task involves a system prompt or tool calls, you need the instruction-following behaviour.
The naming trap on open-weight releases
A single release often ships several variants at once: base, instruct, and sometimes a reasoning or thinking variant. Download listings sort by popularity rather than by which one you want, and quantised community re-uploads frequently drop the suffix entirely.
Two checks before you commit to a download:
Read the model card's intended-use section. A base model card usually says explicitly that the model has not been aligned or instruction-tuned.
Look for a chat template in the repository files. If
tokenizer_config.jsonhas nochat_templatefield, you are almost certainly holding a base model.
The open-weight versus closed model question is a separate axis and the two get conflated constantly. Open weights tell you about the licence and where the model can run. Base versus instruct tells you what the model does when you talk to it.
FAQ
Is an instruct model just a base model with a system prompt?
No. The instruction-following behaviour is in the weights, installed by additional training. A system prompt on a base model does not produce it, because the model has no training that treats a system section as authoritative.
Can I use a base model as a chatbot?
Poorly. You can get some of the way there with few-shot examples and careful stop sequences, but you are recreating by hand what supervised fine-tuning does in the weights, and the result is fragile.
Why does my open-weight model keep talking after answering?
It is most likely a base model, which has no end-of-turn token and no reason to stop. Either switch to the instruct variant of the same release or set an explicit stop sequence.
Does the base model still exist inside an instruct model?
Yes, in the sense that instruction tuning modifies weights rather than replacing them. The knowledge from pretraining is what the instruct model is drawing on. That is also why an instruct model inherits the base model's knowledge cutoff and its factual gaps.
Comparing base and instruct outputs by eye only gets you so far at scale. See what is LLM-as-a-judge for how teams automate that comparison with a second model doing the grading.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


