Dashboard

What Is Pretraining in AI? The Stage Before the Chatbot

Pretraining is the expensive first stage where a model learns to predict the next token. Everything that makes it feel like an assistant happens later.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
20 September 20261 min read

Pretraining is the first and by far the most expensive stage of building a language model. The model reads an enormous amount of text and learns to predict the next token, over and over, until its internal weights encode a usable statistical picture of language and the world described in that text. That is the whole objective. Nobody tells it facts, rules, or manners. It gets one task, next-token prediction, repeated trillions of times.

Everything you associate with a chat assistant, following instructions, refusing things, sounding helpful, comes later. If you want the full pipeline rather than this one stage, start with how AI models work end to end.

What is pretraining in AI, concretely

Take a corpus: web text, books, code, papers. Chop it into tokens. Show the model a sequence, hide the next token, ask it to guess, measure how wrong it was, nudge every weight slightly in the direction that would have been less wrong. Repeat across the corpus for weeks or months on thousands of accelerators.

The result is a base model. It completes text. Ask a base model "What is the capital of France?" and a plausible completion is another question, because in its training data, questions often appear in lists of questions. It is not being difficult. It has no concept that you wanted an answer, because nothing in pretraining taught it that a question implies an obligation to respond.

This is why the shape of the corpus matters so much. The model's ceiling on any subject is roughly what the corpus contained about that subject. It is also why models have a training cutoff date: pretraining happened once, on a fixed snapshot, and the world kept moving.

Pretraining versus everything after it

The stages are easy to confuse because vendors use the words loosely. The practical distinction is what changes and who pays for it.

Stage

What it does

Who runs it

Rough cost

|---|---|---|---|

Pretraining

Builds general capability from raw text

Frontier labs only

Tens to hundreds of millions of dollars

Post-training

Turns a base model into an assistant that follows instructions

Labs, occasionally large enterprises

Meaningfully cheaper than pretraining

Fine-tuning

Adapts an existing model to a narrow task or style

Anyone with an API and a dataset

Tens to thousands of dollars

Prompting

Steers behaviour at request time, changes no weights

You, right now

Cost of tokens

If you want the next stage in detail, post-training is where instruction following and refusals get installed, and fine-tuning is the lever you can actually pull yourself.

The practical implication for builders: you are almost never pretraining. When somebody says they "trained a model on our data," they nearly always mean fine-tuning, and occasionally they mean retrieval, which trains nothing at all.

Why pretraining costs what it does

Three things multiply together. The number of parameters, the number of tokens processed, and the number of times the whole thing is redone after something goes wrong.

The third factor is the one nobody budgets for. A pretraining run can diverge halfway through, produce a loss spike that never recovers, or finish and simply underperform its predecessor. There is no partial credit. You have spent the compute either way.

This is also why frontier labs and small builders are in genuinely different businesses. A lab amortises a nine-figure pretraining run across millions of API customers. Nobody else can run that arithmetic, which is why the open-weight ecosystem depends on a handful of organisations choosing to release the artefact rather than on many organisations producing one.

What pretraining does not fix

A few failure modes get blamed on pretraining that it cannot solve.

  • Recency. The corpus is frozen. Retrieval or web search closes that gap, retraining does not, at least not on any useful timescale.

  • Your private data. It was not in the corpus, and putting it there for the next run is not an option available to you.

  • Reliability on rare cases. Pretraining optimises average next-token accuracy across a corpus. It has no notion of which errors are expensive in your product.

  • Arithmetic and counting. These fall out of a statistical objective badly, which is why they remain a known weak spot long after scale should have solved them.

Frequently asked questions

Is pretraining the same as training?

"Training" is the umbrella term for anything that updates weights. Pretraining is the specific first stage on a general corpus. Fine-tuning and post-training are also training, on much smaller, more targeted data.

Can I pretrain my own model?

For a small model on a narrow domain, yes, and people do. For anything competitive with a current frontier model, no, and the gap is capital, not cleverness. The realistic path for almost every builder is prompting, retrieval, or fine-tuning an existing model.

Does a bigger pretraining corpus always make a better model?

No. Data quality and deduplication matter at least as much as raw volume, and the relationship between compute, data, and capability follows scaling laws rather than simple addition. The Chinchilla paper, Training Compute-Optimal Large Language Models, is the standard reference for why the balance between parameters and tokens matters more than either number alone. Past a point, more of the same low-quality text adds cost without adding capability.

Where does the model learn to refuse things?

Not in pretraining. A base model will happily continue almost any text. Refusals, tone, and instruction following are installed in post-training.

How long does pretraining take?

Weeks to months of wall-clock time on large accelerator clusters for a frontier model, and that is after months of data preparation. The compute is booked in advance, which is part of why release dates slip.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.