What Is an AI Model Card? The Three Fields That Matter
Most of a model card is context you will never use. Three fields on it change what you build, and almost everyone skims past them.
An AI model card is a short structured document that ships alongside a model and states what it was trained to do, what it was evaluated on, and where it is known to fail. Think of it as the spec sheet you would expect for any other component you are about to put in production. Most of a model card is context you will never use. Three fields on it will change what you build, and everyone skims past them.
Where model cards came from
The format comes from a 2018 paper, Model Cards for Model Reporting, which proposed that every trained model ship with documentation covering intended use, out-of-scope use, training data, evaluation results, and known limitations. Hugging Face later made the format a first-class part of its hub, and its model card documentation is now the de facto template for open models.
Commercial labs publish something similar under different names. A system card is the same idea stretched to cover a whole deployed product rather than a set of weights, which is why it usually leads with safety evaluations instead of task benchmarks. If you want the distinction in detail, how to read an AI model's system card covers the longer document.
The three fields that matter
1. Out-of-scope use
This is the section that tells you where the lab expects the model to fail, written by the people with the most information about it. It is also the section most likely to describe exactly what you were planning to do.
A card that says "not evaluated for use in medical, legal, or financial decision-making" is not boilerplate. It means nobody measured performance on that distribution, so your own eval is the only evidence that exists. If your product sits in one of those areas, the out-of-scope list is your test plan.
2. Evaluation data, not evaluation scores
Everyone reads the scores. The useful line is the one above them, naming the dataset. A model reporting 91% on a benchmark built from public web text tells you almost nothing about performance on your customers' scanned PDFs or your internal ticket language.
The question to ask is whether the evaluation distribution resembles your input distribution. When it does not, the number is a comparison tool between models and nothing more. This is the same reasoning behind why a benchmark score can be misleading even when it is honestly reported.
3. Training data cutoff and composition
The cutoff date sets what the model can know without retrieval. The composition, when disclosed, sets what it is fluent in. A model trained mostly on English web text will handle your German support tickets worse than the headline multilingual claim suggests, and the card is where that gets admitted, usually in one sentence.
A worked read
Here is the shape of the read, on any card, in about four minutes.
Field | What you are looking for | What it changes |
|---|---|---|
Intended use | Whether your task is named | Whether you need your own eval before shipping |
Out-of-scope use | Whether your task is excluded | Whether you need a human review step |
Training data | Cutoff date and language mix | Whether you need retrieval or translation |
Evaluation data | Distribution similarity to yours | How much to trust the reported scores |
Limitations | Named failure modes | What your monitoring should watch for |
Licence | Commercial use, redistribution, fine-tune rights | Whether you can legally ship it at all |
The licence row is the one that ends projects. Weights described as "open" carry a wide range of terms, some of which restrict commercial use or cap monthly active users, which is why AI model licences deserve their own read rather than a glance at a badge.
What a model card does not tell you
It does not tell you how the model behaves on your data. No card can. It is a starting point that narrows what you need to test, not a substitute for testing.
It also does not stay true. Providers update models behind the same name, and a card written for the first release may not describe the version answering your calls today. This is one of the reasons testing a new model before switching is worth the afternoon it costs, and why a small held-out eval set of your own is more durable than any published document.
Finally, a card is written by the party selling the model. Limitations sections are honest more often than cynics expect, but they are scoped by what the lab chose to measure. Absence of a known failure mode is not evidence of its absence, and the broader picture of how AI models work will tell you more about likely failure shapes than any single card.
FAQ
What is the difference between a model card and a system card?
A model card documents a set of weights: task, training data, evaluations, limitations. A system card documents a deployed product built around a model, including safety mitigations, and usually covers policy behaviour the weights alone do not determine.
Do closed commercial models have model cards?
Most publish something equivalent, often as a documentation page rather than a file. The fields are usually there even when the format is not: context window, knowledge cutoff, intended use, and a limitations section.
Is a model card legally binding?
No. It is documentation, not a warranty. Your contractual position comes from the provider's terms of service and the model licence, not the card.
How long should reading one take?
Four to five minutes if you go straight to intended use, out-of-scope use, training cutoff, evaluation data and licence. Reading it end to end is rarely worth it unless you are choosing between two close options.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


