Model names now carry two numbers: Qwen3-30B-A3B, gpt-oss-20b with "20.9B total, 3.6B active". Those are mixture-of-experts (MoE) models, and the two numbers answer different questions. Total parameters decide how much memory you need. Active parameters decide how fast it runs. Mixing them up leads to the two most common mistakes: buying hardware for the wrong number, and expecting a 30B model's quality from something that computes like a 3B one.

How an MoE layer works

In a standard ("dense") transformer, every token passes through every weight. In an MoE model, the big feed-forward part of each layer is split into many experts, and a small router picks a few of them for each token. The rest of that layer's experts sit idle for that token.

The published numbers make the ratio concrete:

Model Total parameters Active per token Experts
Mixtral 8x7B 46.7B 12.9B 8 per layer, 2 chosen
Qwen3-30B-A3B 30.5B 3.3B 128 per layer, 8 chosen
gpt-oss-20b 20.9B 3.6B 32 per layer, 4 chosen

Sources: Mistral's Mixtral announcement, the Qwen3-30B-A3B model card, the gpt-oss model card.

Attention layers and embeddings are shared and always active. Only the expert blocks are sparse.

Memory follows the total

The router can choose any expert for the next token, so every expert has to be available. A 30B MoE model needs roughly the memory of a 30B dense model at the same quantization. Quantization helps both equally (see my piece on Q4 and Q8); sparsity doesn't reduce what you store.

Speed follows the active count

Generating one token reads the shared weights plus the chosen experts. For single-request generation, which is bound by memory bandwidth, that's the number that matters. Mistral put it plainly when it released Mixtral: the model processes input and generates output at the speed and cost of a 12.9B model while having 46.7B parameters in total (announcement).

Batching muddies this on servers. When many requests are processed together, different tokens route to different experts, and together they touch most of the model. The speed advantage per token is largest for one user on one machine, which is exactly the local case.

Quality sits between the two numbers

An MoE model's capacity comes from all of its parameters: different experts specialize during training. Its per-token computation is small. In practice a well-trained MoE model tends to perform better than a dense model of its active size and not as well as a dense model of its total size. Treat that as a tendency and compare specific models on your task.

The local trick: experts on the CPU

Because only a few experts run per token, you can keep the expert weights in system RAM and everything else on the GPU. llama.cpp supports this directly: --n-cpu-moe N keeps the expert tensors of the first N layers on the CPU, while attention, shared weights and the KV cache stay on the GPU; -ot (override-tensor) gives finer control with regular expressions over tensor names (a detailed guide on Hugging Face).

On a small GPU this is the difference between not running a model and running it at a usable speed. The CPU does the expert multiplications for a few experts per layer, reading from RAM; the GPU handles attention with the context in fast memory. Raise N until the model fits with the context you need, then measure tokens per second.

My local pipeline does most of its language work with MoE models such as gpt-oss-20b on an 8 GB laptop GPU, and splitting weights between GPU and system RAM is what makes that workable. A dense model of the same total size gains much less from the same split, because every token needs every weight, including the ones in slow RAM.

When to choose MoE

  • You're memory-rich and compute-poor: plenty of RAM, a small GPU or none. MoE with expert offload is often the best quality you can run.
  • You serve one user or a few at a time, where the active-parameter speedup shows up fully.
  • You want a large model's breadth of knowledge at small-model latency.

A dense model is the better fit when memory is the tight constraint (a dense model of the MoE's active size fits in much less memory), when you plan to fine-tune (MoE fine-tuning is more involved), or when your serving setup batches heavily and the per-token advantage shrinks.

Reading a model card

When you see a model with "A" in its name or two parameter counts, write down both numbers, the number of experts and how many are chosen per token. Size your memory for the total, estimate speed from the active count, and then test quality on your own inputs, because neither number predicts that on its own.

The practical next step

Map one real execution and one failure.

That will reveal more about the right architecture than a tool comparison or model demo.

Let's build something real