What Is Pretraining in AI? The Stage Before the Chatbot
Pretraining is the expensive first stage where a model learns to predict the next token. Everything that makes it feel like an assistant happens later.
Pretraining is the first and by far the most expensive stage of building a language model. The model reads an enormous amount of text and learns to predict the next token, over and over, until its internal weights encode a usable statistical picture of language and the world described in that text. That is the whole objective. Nobody tells it facts, rules, or manners. It gets one task, next-token prediction, repeated trillions of times.
Everything you associate with a chat assistant, following instructions, refusing things, sounding helpful, comes later. If you want the full pipeline rather than this one stage, start with how AI models work end to end.
What is pretraining in AI, concretely
Take a corpus: web text, books, code, papers. Chop it into tokens. Show the model a sequence, hide the next token, ask it to guess, measure how wrong it was, nudge every weight slightly in the direction that would have been less wrong. Repeat across the corpus for weeks or months on thousands of accelerators.
The result is a base model. It completes text. Ask a base model "What is the capital of France?" and a plausible completion is another question, because in its training data, questions often appear in lists of questions. It is not being difficult. It has no concept that you wanted an answer, because nothing in pretraining taught it that a question implies an obligation to respond.
This is why the shape of the corpus matters so much. The model's ceiling on any subject is roughly what the corpus contained about that subject. It is also why models have a training cutoff date: pretraining happened once, on a fixed snapshot, and the world kept moving.
Pretraining versus everything after it
The stages are easy to confuse because vendors use the words loosely. The practical distinction is what changes and who pays for it.
Stage | What it does | Who runs it | Rough cost |
|---|
|---|---|---|---|
Pretraining | Builds general capability from raw text | Frontier labs only | Tens to hundreds of millions of dollars |
|---|---|---|---|
Post-training | Turns a base model into an assistant that follows instructions | Labs, occasionally large enterprises | Meaningfully cheaper than pretraining |
Fine-tuning | Adapts an existing model to a narrow task or style | Anyone with an API and a dataset | Tens to thousands of dollars |
Prompting | Steers behaviour at request time, changes no weights | You, right now | Cost of tokens |
If you want the next stage in detail, post-training is where instruction following and refusals get installed, and fine-tuning is the lever you can actually pull yourself.
The practical implication for builders: you are almost never pretraining. When somebody says they "trained a model on our data," they nearly always mean fine-tuning, and occasionally they mean retrieval, which trains nothing at all.
Why pretraining costs what it does
Three things multiply together. The number of parameters, the number of tokens processed, and the number of times the whole thing is redone after something goes wrong.
The third factor is the one nobody budgets for. A pretraining run can diverge halfway through, produce a loss spike that never recovers, or finish and simply underperform its predecessor. There is no partial credit. You have spent the compute either way.
This is also why frontier labs and small builders are in genuinely different businesses. A lab amortises a nine-figure pretraining run across millions of API customers. Nobody else can run that arithmetic, which is why the open-weight ecosystem depends on a handful of organisations choosing to release the artefact rather than on many organisations producing one.
What pretraining does not fix
A few failure modes get blamed on pretraining that it cannot solve.
Recency. The corpus is frozen. Retrieval or web search closes that gap, retraining does not, at least not on any useful timescale.
Your private data. It was not in the corpus, and putting it there for the next run is not an option available to you.
Reliability on rare cases. Pretraining optimises average next-token accuracy across a corpus. It has no notion of which errors are expensive in your product.
Arithmetic and counting. These fall out of a statistical objective badly, which is why they remain a known weak spot long after scale should have solved them.
Frequently asked questions
Is pretraining the same as training?
"Training" is the umbrella term for anything that updates weights. Pretraining is the specific first stage on a general corpus. Fine-tuning and post-training are also training, on much smaller, more targeted data.
Can I pretrain my own model?
For a small model on a narrow domain, yes, and people do. For anything competitive with a current frontier model, no, and the gap is capital, not cleverness. The realistic path for almost every builder is prompting, retrieval, or fine-tuning an existing model.
Does a bigger pretraining corpus always make a better model?
No. Data quality and deduplication matter at least as much as raw volume, and the relationship between compute, data, and capability follows scaling laws rather than simple addition. The Chinchilla paper, Training Compute-Optimal Large Language Models, is the standard reference for why the balance between parameters and tokens matters more than either number alone. Past a point, more of the same low-quality text adds cost without adding capability.
Where does the model learn to refuse things?
Not in pretraining. A base model will happily continue almost any text. Refusals, tone, and instruction following are installed in post-training.
How long does pretraining take?
Weeks to months of wall-clock time on large accelerator clusters for a frontier model, and that is after months of data preparation. The compute is booked in advance, which is part of why release dates slip.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


