Dashboard

What Is Model Merging in AI? A Clear Explanation

Model merging combines the weights of two or more fine-tuned models into a single model, no training involved. Here is how linear averaging, SLERP, and TIES merging actually differ.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
5 September 20261 min read

Model merging is the practice of combining the weights of two or more fine-tuned models into a single model, without any additional training. You take models that started from the same base, each fine-tuned for a different skill, and blend their parameters directly. The result is one model that inherits traits from all of them, at the speed of running just one model.

This is not the same as fine-tuning and not the same as distillation, even though people mix the three up. Getting the difference right matters, because the wrong technique wastes GPU hours on a problem arithmetic could have solved in minutes.

Model merging vs. fine-tuning vs. distillation

These three terms describe different operations on different numbers of models, and confusing them leads to the wrong project plan.

The common thread is that all three operate on models built from the foundation model that every fine-tuned variant descends from. Merging in particular only works cleanly when every input model shares that same base and parameter shapes. You cannot merge a Llama-based checkpoint with a Mistral-based one, because the weights do not correspond to the same positions in the same structure.

Why merge models at all

Say you fine-tuned one checkpoint to be better at code generation, and a separate checkpoint of the same base to be better at following multi-step instructions. Running both in production means double the inference cost and picking which one answers each request. Merging combines both sets of learned changes into one checkpoint that tries to do both, at the cost of running a single model. It also recovers skills a later fine-tuning pass eroded, and lets you combine community checkpoints without owning the original training data. None of that touches the training pipeline, it just requires picking a method that matches what you want to preserve.

Three ways to actually merge models

There is no single merging algorithm. The three below range from simplest to most careful, and each assumes something different about how the models' changes interact.

1. Linear averaging and task arithmetic model merging

The simplest merge, often called model soup, just averages the raw weights of two or more fine-tuned models position by position. It works surprisingly well when the models were fine-tuned from the same starting checkpoint on similar tasks with similar hyperparameters.

Task arithmetic model merging refines this by working with task vectors instead of raw weights, the difference between a fine-tuned model and its base, that is, everything fine-tuning actually changed. It computes a task vector for each fine-tune, optionally scales each one, sums them, and adds the result back onto the base model. Because it operates on deltas rather than absolute weights, you can also subtract a task vector to suppress a behavior, not just add one to gain it.

Use this when the models you are merging make small, non-conflicting changes to the base, such as combining a formatting fine-tune with a tone fine-tune. It is the cheapest method to run and the right default to try first.

2. SLERP model merging

SLERP, spherical linear interpolation, treats each model's weights as a point on a high-dimensional sphere rather than a point in flat space. Instead of drawing a straight line between two models and averaging along it, SLERP follows the curved, shortest arc between them on that sphere. That preserves the geometric relationships and relative magnitudes within each model's weights better than a straight average, instead of just blurring both models toward the middle.

SLERP is defined for two models at a time. Tools such as the mergekit library extend it further through multi-step interpolation, but the core operation is pairwise. It tends to produce smoother, more coherent merges than linear averaging when the two source models diverged meaningfully from the base, exactly the case where flat averaging starts to wash out both models' distinct behavior.

3. TIES merging

TIES merging solves a specific failure mode: when you merge three or more task vectors, some of their parameter changes point in opposite directions and cancel out, or a handful of large, noisy changes drown out everything else. TIES, short for Trim, Elect Sign, and Disjoint Merge, fixes this in three steps. Trim discards each task vector's smallest, most likely noisy values. Elect Sign looks at every parameter where the remaining vectors disagree on direction and keeps whichever direction has the larger total magnitude across models. Disjoint Merge then averages only the vectors agreeing with that elected sign, instead of blending all of them regardless of conflict.

This makes TIES the better choice when merging several models fine-tuned on genuinely different tasks, where sign conflicts are likely. A related method, DARE, takes a different route to the same problem, randomly dropping most of each task vector's values before rescaling the rest.

Picking a method for your merge

  • Two models, similar tasks, small changes from the base: start with linear averaging or task arithmetic, the fastest baseline to compare everything else against.

  • Two models that diverged a lot from the base, where a straight average degrades both: try SLERP before reaching for anything more complex.

  • Three or more models fine-tuned on genuinely different tasks: use TIES, or DARE-TIES, to handle sign conflicts a plain average would let cancel out.

Evaluate the merged model on held-out tasks from each source model before shipping it. Merging is cheap enough to try more than one method and compare, which is the point: it lets you experiment with combinations that would each cost a full training run as fine-tunes. For the broader picture of what a model's weights represent and why combining them works, see how the weights inside an AI model actually work.

Frequently asked questions

Is model merging the same as ensembling?

No. An ensemble keeps every model separate and runs all of them at inference time, then combines their outputs, which multiplies compute and latency. Merging combines the weights ahead of time into one model, so you deploy and run just one at normal cost.

Do merged models need any retraining?

Not for the merge itself. Linear averaging, SLERP, and TIES are arithmetic on existing weights and run in minutes on a CPU. Some teams do a short fine-tuning pass afterward to smooth rough edges, but it is optional.

Can you merge models with different architectures?

Generally no. These methods require the models to share the same architecture and parameter shapes, typically fine-tuned checkpoints of the same base model. Merging a Llama-based model with a Mistral-based one is not directly supported.

What is a task vector?

A task vector is the difference between a fine-tuned model's weights and the base model's weights, everything fine-tuning changed. Task arithmetic and TIES operate on task vectors rather than raw weights, which lets them combine several fine-tunes without the base knowledge canceling out.

How many models can you merge at once?

Linear averaging and task arithmetic scale to as many models as you feed in. SLERP is defined for two at a time, though tools like mergekit extend it further. TIES was designed specifically for merging several task-specific models together.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.