Dashboard

Dense Model vs Mixture of Experts: Which Should You Pick?

Dense models run every parameter on every token. Mixture of experts models activate only a handful. Here is what that actually changes for cost, latency, and which to pick as an API caller.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
16 September 20261 min read

Dense Model vs Mixture of Experts: Which Should You Pick?

A dense model runs every one of its parameters on every token it processes. A mixture of experts (MoE) model splits its parameters into specialized sub-networks, or experts, and a router activates only a handful of them per token, so a model with a huge total parameter count can run at the inference cost of a much smaller one. Neither is universally better, they trade off differently on cost, speed, and consistency, and as someone choosing between AI providers and model tiers rather than training your own, the practical question is which tradeoff fits what you are building.

What actually changes between the two

Dense

Mixture of experts

Parameters used per token

All of them

A small subset, chosen by a router

Inference cost relative to total size

Scales with total parameters

Scales with active parameters, much lower

Memory footprint to serve

Roughly matches total parameters

Must hold all experts in memory, even unused ones

Output consistency

Very consistent, same path every time

Can vary slightly based on which experts a router picks

Typical use

Smaller models, latency-critical paths

Frontier-scale models needing huge capacity at usable cost

Why MoE exists at all

Scaling a dense model's quality means scaling every parameter's cost right alongside it, doubling total parameters roughly doubles inference cost per token. Mixture of experts breaks that link: you can grow total capacity by adding more experts without growing the cost of a single forward pass by nearly as much, since the router only activates a few of them. This is why most of today's largest frontier models use an MoE architecture, it is close to the only practical way to keep scaling capacity without scaling latency and cost at the same rate.

The catch: memory does not get the same discount

Inference cost drops because only a few experts run per token, but every expert still has to be loaded into memory somewhere in case the router picks it, because which experts get picked can change token to token, even within the same request. A mixture of experts model with a huge total parameter count needs enough memory (or enough GPUs) to hold the whole thing, even though any single token only uses a fraction of it. This is the real reason MoE models are mostly something specialized providers serve at scale, rather than something you would self-host casually, the memory bill does not shrink the way the compute bill does.

What this means for you as an API caller, not a trainer

If you are building with AI rather than training models, you are choosing between finished models exposed through an API, so the dense-versus-MoE question shows up indirectly, as differences in price, latency, and consistency between model tiers from the same provider, more than as an explicit architecture toggle.

  • If a task needs the largest available reasoning capability and cost per token is secondary, a frontier MoE model is usually the strongest option, that architecture is how providers reach the highest capability tier at all.

  • If a task is latency-sensitive, high-volume, or running inside a tight per-request budget (autocomplete, a chat widget handling many concurrent users, a background classification job), a smaller dense model is often the better fit, its cost and latency are more predictable and it avoids router-related variance.

  • If you notice a provider's flagship model occasionally gives a subtly different style of answer to near-identical prompts, that is sometimes the router sending similar-but-not-identical requests to different experts, a known MoE characteristic, not a bug in your prompt.

A rule of thumb for picking between model tiers

Match the model to the task's actual difficulty rather than defaulting to the largest, most expensive tier for everything. A frontier MoE model is worth its cost on genuinely hard reasoning, planning, or code-generation tasks. For high-volume, lower-complexity work, a smaller dense model frequently performs the task just as well at a fraction of the latency and cost, run a side-by-side comparison on your actual use case rather than assuming bigger is always better for what you specifically need.

Where this connects

The router mechanism itself, and how it decides which experts to activate, is worth understanding on its own, see what mixture of experts is for the deeper mechanical explanation. For the broader vocabulary of model size and cost, see what a parameter actually is and how quantization shrinks a model further, a separate lever from the dense-versus-MoE choice that can be applied to either architecture.

For the full builder's map of how models work underneath the API you call, start at how AI models work.

Frequently asked questions

Are all the largest AI models mixture of experts now?

Most frontier-scale models from major labs use some form of MoE architecture today, since it is the main lever for growing total capacity without a proportional growth in inference cost. Smaller, more specialized, or latency-focused models are still frequently dense.

Does mixture of experts mean the model is less capable per parameter?

Not in a way that matters practically. The point of MoE is reaching a useful capability level at a lower inference cost than a dense model of equivalent total size would need, not sacrificing quality, the specialization across experts is part of what lets it work well.

Can I tell from a provider's pricing page whether a model is dense or MoE?

Rarely directly, providers do not usually label it on a pricing page. You can infer it indirectly: a model priced far below what its stated capability would suggest for a dense architecture of similar strength is very likely MoE under the hood.

Does this affect which model I should fine-tune?

If you are fine-tuning rather than calling an API, dense models are generally simpler to fine-tune with standard tooling. Fine-tuning a mixture of experts model correctly, including which experts get updated, is a more specialized process, and most builders working with AI instead of training it will not need to worry about this distinction at all.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.