You fine-tuned one model for coding, someone else fine-tuned another for math, and a third checkpoint is great at following instructions. What if you could combine all three into a single model — without another training run, without new data, and without a GPU cluster? That is exactly what model merging does. In 2026 it has become one of the cheapest ways for developers to build capable open-weight models, and a toolkit called MergeKit makes it a matter of writing a short YAML file.
This guide explains how model merging works, compares the main methods (SLERP, TIES, DARE and friends), and walks you through your first merge step by step.
What Is Model Merging?

Model merging combines the weight tensors of two or more trained models into one new set of weights. Unlike ensembling, you do not run several models at inference time — you get a single model with the same size and speed as its parents. Unlike fine-tuning, there is no gradient descent: the merge is pure arithmetic on parameters.
The catch is that the parent models must be compatible. In practice that means they share the same architecture and, ideally, were fine-tuned from the same base model (for example, several Llama or Qwen fine-tunes of one base checkpoint). Merging two unrelated models of different families usually produces garbage.
- No training data needed — useful when the original fine-tuning data is private or lost.
- Runs on modest hardware — MergeKit can merge on CPU or a single GPU with lazy tensor loading.
- Fast iteration — a merge takes minutes to an hour, so you can try many recipes cheaply.
- Zero inference overhead — the output is a normal model you can quantize and serve like any other.
How Model Merging Works: Task Vectors in Plain English
Most modern merge methods build on the idea of a task vector, introduced in the task arithmetic research by Ilharco et al. Take a fine-tuned model and subtract the base model’s weights. What is left — the delta — represents what the model learned during fine-tuning. If you add several of these deltas back onto the base model, you can (with care) stack their skills.
The problem is interference. Two task vectors may push the same parameter in opposite directions, or carry lots of small, noisy changes that cancel each other’s useful signal. Every advanced method below is essentially a different strategy for reducing that interference.
The Main Merge Methods Compared
Linear (Model Soups)
The simplest approach: a weighted average of the models’ weights. It works surprisingly well when the models are close to each other — for example, several fine-tunes of the same base with different hyperparameters — and is a sensible baseline.
SLERP (Spherical Linear Interpolation)
SLERP interpolates between two models along a curved path rather than a straight line, which better preserves the geometry and magnitude of the weights. It is a community favorite for blending two models smoothly, but it only handles two models at a time. You can set different interpolation factors for attention and MLP layers.
Task Arithmetic
Compute each model’s task vector, scale it by a weight, and add the sum to the base. It is easy to reason about and supports many models, but it is vulnerable to interference when you merge more than two or three.
TIES-Merging
TIES (Trim, Elect Sign and Merge) from Yadav et al. tackles interference directly. It trims each task vector to its largest-magnitude changes (controlled by a density parameter), elects a dominant sign for each parameter across models, and then averages only the values that agree with that sign. It is one of the most reliable choices for merging three or more models.
DARE (Drop and Rescale)
DARE, from Yu et al., randomly drops a large fraction of each task vector’s delta parameters and rescales the survivors to compensate. The surprising finding is that fine-tuning deltas are highly redundant, so you can drop most of them with little loss. In MergeKit it is usually combined with TIES sign election as dare_ties, or with plain task arithmetic as dare_linear.
Model Stock, DELLA and Others
MergeKit also supports newer methods such as Model Stock (which uses the geometry of several fine-tunes to choose interpolation weights automatically), DELLA (magnitude-aware dropping), Model Breadcrumbs and Passthrough, which stacks layers from different models to build “frankenmerges” with more layers than either parent.
| Method | Models | Needs base model | Best for |
|---|---|---|---|
| Linear | 2+ | No | Averaging close fine-tunes |
| SLERP | Exactly 2 | No | Smoothly blending two models |
| Task Arithmetic | 2+ | Yes | Simple, controllable skill stacking |
| TIES | 3+ | Yes | Reducing interference in multi-model merges |
| DARE (TIES/linear) | 2+ | Yes | Sparse, redundant fine-tune deltas |
| Model Stock | 3+ | Yes | Automatic weighting of many fine-tunes |

Step-by-Step: Your First Merge with MergeKit
MergeKit is an open-source toolkit from Arcee AI, described in detail in the MergeKit paper on arXiv. It supports Llama, Mistral, Qwen, Gemma and many other architectures. Check the repository for its current license terms before using it commercially.
1. Install MergeKit
git clone https://github.com/arcee-ai/mergekit.git
cd mergekit
pip install -e .2. Write a merge config. Here is a TIES merge of two fine-tunes that share one base model (replace the model names with your own):
merge_method: ties
base_model: your-org/base-7b
models:
- model: your-org/math-finetune-7b
parameters:
density: 0.5
weight: 0.5
- model: your-org/code-finetune-7b
parameters:
density: 0.5
weight: 0.5
parameters:
normalize: true
dtype: bfloat163. Run the merge
mergekit-yaml config.yml ./merged-model --cuda --lazy-unpickleDrop --cuda to run on CPU. The output folder is a standard Hugging Face model directory that you can load with Transformers, push to the Hub, or quantize for local inference.
4. Evaluate — always. Merges can look great on one benchmark and quietly break chat formatting or safety behavior. Run your own task-specific evals before you ship. Our guide to LLM evaluation with Ragas, DeepEval and Promptfoo covers practical tooling.
Model Merging Best Practices
- Start from a shared base. Merges of siblings from one base checkpoint are far more stable than cross-lineage merges.
- Match chat templates and tokenizers. Mismatched special tokens are a common cause of broken output; MergeKit has options to control how the tokenizer is built.
- Tune density and weight gradually. For TIES and DARE, densities around 0.3–0.7 are typical starting points — sweep a few values rather than guessing.
- Prefer SLERP for two models, TIES or DARE for three or more.
- Re-check safety. Merging can dilute alignment learned by one parent. Red-team the result just as you would a fresh fine-tune.
- Quantize after merging, not before. Merge in bf16 or fp16, then apply GGUF or AWQ — see our LLM quantization guide.
When Model Merging Is the Wrong Tool
Merging only recombines what the parents already know. If no parent has a skill, the merge will not either. If you need a model to learn new knowledge or a new output format, fine-tune it — our comparison of LoRA vs QLoRA vs full fine-tuning and our Unsloth vs Axolotl vs LLaMA-Factory guide will help. If you need a smaller, faster model, look at knowledge distillation instead. A popular hybrid workflow is to train several LoRA adapters on different tasks, merge each into the base, and then combine the results with TIES.

Frequently Asked Questions
Does model merging require a GPU?
No. MergeKit can run entirely on CPU with lazy loading, though a GPU speeds things up considerably. Memory, not compute, is usually the main constraint for large models.
Can I merge models from different families, like Llama and Mistral?
Not with standard weight-averaging methods. Models must share the same architecture and tensor shapes, and results are best when they were fine-tuned from the same base checkpoint.
What is the difference between TIES and DARE?
TIES trims small changes and resolves sign conflicts between models. DARE randomly drops most delta parameters and rescales the rest. They are complementary, which is why dare_ties combines both.
Is a merged model better than a fine-tuned one?
Sometimes. A good merge can combine strengths from several fine-tunes and score above each parent on broad benchmarks, but it cannot add skills that none of the parents have. Always evaluate on your own tasks.
Conclusion
Model merging lets you combine the strengths of several fine-tuned LLMs into one model in minutes, with no training data and modest hardware. Start with SLERP for two models, move to TIES or DARE when you merge more, and treat evaluation as non-negotiable. Ready to try it? Pick two fine-tunes of the same base, run the MergeKit config above, and compare the result against both parents — then explore more of our AI and LLM guides to take your merged model into production.

