Download an open-weight model for a laptop or a small server and you'll face a list of files: Q4_K_M, Q5_K_M, Q6_K, Q8_0, F16. The labels look like noise, but they answer three practical questions: will it fit, how fast will it run, and how much worse will it be. This is the version of the explanation I wish I'd had when I started running models on an 8 GB GPU.
What quantization does
A model is mostly weights: billions of numbers, trained in 16- or 32-bit floating point. Quantization stores each weight with fewer bits. Instead of a 16-bit float, a block of weights shares one scale factor and each weight becomes a small integer (8 bits, 5, 4, sometimes fewer). At inference time the weights are expanded back just in time for the multiplication.
The naming in llama.cpp's GGUF format encodes the scheme:
- The number is roughly the bits per weight: Q4 is about four, Q8 about eight.
Kmarks the "k-quant" family, which quantizes in blocks with extra scale information and generally keeps more quality per bit than the older formats.S,MandL(small, medium, large) say how many sensitive tensors get extra precision.Q4_K_Mkeeps some layers at higher precision thanQ4_K_S._0formats such asQ8_0are simpler block formats; at 8 bits, simple is plenty.
The effective bits per weight are a little higher than the name, because scales cost space too. llama.cpp's own table for Llama-3.1-8B lists Q4_K_M at about 4.89 bits per weight and Q8_0 at about 8.50 (quantize README).
Question one: will it fit?
The memory for weights is close to parameters × bits per weight ÷ 8:
8B params × 4.89 bits ÷ 8 ≈ 4.9 GB (the README lists 4.58 GiB for Q4_K_M)
8B params × 8.50 bits ÷ 8 ≈ 8.5 GB (7.95 GiB for Q8_0)
8B params × 16 bits ÷ 8 ≈ 16 GB (14.96 GiB for F16)
(The README reports GiB, and 1 GiB is about 1.07 GB, which accounts for most of the gap.)
Weights aren't the whole bill. The KV cache, which stores attention state for every token in the context, grows with context length and can take gigabytes on its own at long contexts. Leave room for it. On my 8 GB card, a 12B model at Q4 fits with a moderate context; the same model at Q8 doesn't.
Question two: how fast?
When a model generates text one token at a time on a single request, the GPU spends most of its time reading weights from memory, not computing. Fewer bytes per weight means fewer bytes to read, so generation gets faster as the file gets smaller. The same README shows this on its test hardware: text generation for Llama-3.1-8B ran at about 29 tokens per second in F16, 51 in Q8_0 and 72 in Q4_K_M.
Prompt processing is a different workload. It handles many tokens at once and is limited by compute, so it barely benefits from quantization and can even run slower on some formats. That table shows prompt processing slightly faster in F16 than in the quantized versions. If your workload is long documents in and short answers out, measure both numbers.
Question three: how much worse is Q4_K_M than Q8_0?
This is the question people answer with vibes. llama.cpp publishes a better answer: compare each quantized model's token probabilities with the full-precision model's on the same text. Two numbers matter. KL divergence measures how far the probability distributions moved (lower is better). "Same top token" is how often both models would pick the same next token.
For Llama 3 8B, the project's scoreboard reports (perplexity README):
| Format | Size | Mean KL divergence | Same top token |
|---|---|---|---|
| Q8_0 | 7.96 GiB | 0.0014 | 97.7% |
| Q5_K_M | 5.33 GiB | 0.0108 | 96.0% |
| Q4_K_M (with an importance matrix) | 4.58 GiB | 0.0282 | 91.9% |
Q8_0 is very close to the original. If you run mixture-of-experts models, the same memory arithmetic applies to their total parameter count (more on MoE memory and speed). Q4_K_M disagrees about the top token roughly once in twelve tokens. Most of those disagreements are between two plausible words and change nothing. Some aren't, and they compound over a long answer.
Averages hide where the damage lands. Quantization tends to hurt more on the things a model was already bad at: rare languages, exact numbers and long chains of reasoning. Smaller models also suffer more than larger ones at the same bit width, because they have less redundancy to lose. A 70B model at Q4 is usually a better deal than an 8B model at Q8; a 3B model at Q4 can fall apart on tasks it handled at Q8.
Importance matrices and other refinements
An importance matrix (imatrix) runs sample text through the model to find which weights matter most, and spends precision there. The Q4_K_M row above used one. Other formats exist: GPTQ and AWQ for GPU servers, and MXFP4, a 4-bit floating-point format that OpenAI used for the mixture-of-experts weights of its gpt-oss models (gpt-oss model card). The same questions apply to every format.
How I pick
- Start from the memory budget: weights plus the KV cache for the context you actually need.
- Prefer the largest model that fits at Q4_K_M or Q5_K_M over a smaller model at Q8.
- Use Q8_0 when the model is small, the task needs exact numbers or code, or memory isn't the constraint.
- Test on my task before believing any of the above.
Test it on your task
Perplexity tables are general. Your task isn't. Take 50 to 100 real inputs with known good outputs and run them through two quantizations of the same model at the same settings. Score them the way you'd score production output: exact match for extraction, a checklist for summaries, a blind human preference for writing. Record tokens per second for both prompt processing and generation on your hardware.
If the difference isn't visible on your own cases, take the smaller file and spend the saved memory on context or a bigger model. If it is, you've just learned something no public table could tell you.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real