Dashboard

Should You Pin an AI Model Version?

An unpinned model alias changes underneath you on the vendor's schedule. When to pin, what pinning costs, and the config setup that gets you both.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
19 September 20261 min read

Should You Pin an AI Model Version?

Pin in production, float in development. That is the short answer, and the reason is that an unpinned alias changes underneath you at a time the vendor chooses, which is fine while you are experimenting and unacceptable while customers are using the output. The longer answer is about what pinning costs you, because it is not free.

The case for pinning

Pinning buys you one thing: the output distribution stops moving on its own. That matters more than it sounds, because prompts are tuned against a specific model's habits. A prompt that reliably returns clean JSON from one snapshot may start wrapping it in explanation on the next, and nothing in your code changed.

  • Regressions become attributable. If output quality drops, it was your change, not a silent swap.

  • Evals stay comparable over time. An unpinned baseline makes month-over-month scores meaningless.

  • Incidents are reproducible. You can re-run last Tuesday's failing request and get last Tuesday's behaviour.

  • Cost stays predictable. A newer snapshot in the same family can have different pricing or different verbosity.

Knowing which one you are using

The distinction is not always obvious from the identifier string itself, which is a recurring source of accidents. Most providers expose both styles for the same family:

Identifier style

Behaviour

Where it belongs

Floating alias, e.g. a bare family name

Points at whatever the vendor currently considers current; changes without your involvement

Local development, evals, exploration

Dated or numbered snapshot

Fixed weights, fixed behaviour, until the vendor retires it

Production, anything with a prompt you tuned

What the version numbers actually mean goes through how the different vendors name things, and it is worth checking rather than assuming, because a string that looks specific is sometimes still an alias.

The case against, which is real

Pinning is not a safe default you can set and forget. It creates two specific debts.

The first is that pinned snapshots get retired. A pin is a commitment to migrate on the vendor's schedule, and if you pinned and stopped paying attention, the retirement notice is the first you hear of it. That is a worse position than floating, because floating would have moved you gradually. Handling a deprecation when it lands covers the scramble, and the way to avoid it is to treat a pin as a dated obligation rather than a setting.

The second is that you stop getting improvements. Newer snapshots are usually better at instruction following and often cheaper per token. A team that pinned eighteen months ago and never revisited is paying more for worse output and calling it stability.

The setup that gets both

Pin the model in one place, not scattered through the codebase, and make the pin an environment variable rather than a literal. Then the upgrade is a config change you can roll back, not a deploy.

bash
# .env.production
MODEL_ID=<dated-snapshot-id>

# .env.development
MODEL_ID=<floating-alias>

With that in place, the routine is: development runs against the floating alias, so you notice changes early and cheaply. Before promoting a new snapshot to production, run your eval set against both and compare. Promote by changing one variable.

How often to revisit

  1. Put a recurring calendar reminder, monthly or quarterly, to check the provider's deprecation page for your pinned snapshot.

  2. Keep a small eval set, twenty prompts is plenty, with expected-output notes rather than exact matches.

  3. When a new snapshot appears, run the set against both and read the diffs rather than a score.

  4. Promote when the new one is not worse. You do not need it to be dramatically better to take the cheaper, longer-supported option.

One thing that eval set protects you from: the very common experience of a model apparently degrading after an update, which is sometimes real and sometimes a shift in output style that your prompt was quietly relying on. Why a model can seem worse after an update pulls those two apart, and having recorded outputs from before the change is the only way to tell which you are looking at.

For models that have been announced but not shipped, see how to plan around an unreleased AI model.

Frequently asked questions

Does pinning cost more?

Not directly, but an older snapshot can be more expensive per unit of useful output than a newer one, and older families sometimes keep older pricing. Compare on cost per successful result, not cost per token.

What happens when a pinned version is retired?

The provider publishes a retirement date and, after it, requests to that identifier fail. This is why a pin needs a calendar reminder attached to it.

Should I pin in a side project?

Not really. The overhead is not worth it below the point where somebody else depends on the output being stable.

Is pinning enough to make output deterministic?

No. Sampling settings still introduce variation, and providers can change serving infrastructure. Pinning removes the largest source of drift, not all of it. The mechanics are in how AI models work underneath.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.