Quantization
Quantization stores a model's weights in fewer bits, which cuts their memory roughly in proportion; it makes inference faster only where your runtime and GPU have kernels for that format, and how much quality you lose depends on the method, not just the bit count.
Same model, different representation
A model's weights are a few billion numbers. Published checkpoints usually store each one as BF16, a 16-bit float, which is 2 bytes per parameter. A quantized build stores the same numbers in fewer bits and adds scale factors that map the small values back to roughly the originals. The architecture, tokenizer and API stay the same. The file changes, the memory it takes changes, and the runtime needs code that can read the new format.
Three terms get mixed up:
- Precision is the number format one value is stored in: BF16, FP8, FP4, INT4. Quantization is converting a model from a higher precision to a lower one, with the scale factors that make it work.
- The container is the file format: safetensors or GGUF. A GGUF file can hold 16-bit weights, and a safetensors file can hold 4-bit ones.
- The method is how the numbers are rounded, grouped and scaled: FP8, AWQ, GPTQ, NVFP4, Q4_K_M.
This guide covers the weights. The KV cache (the memory each token of context takes) can be quantized separately. See Context windows and KV cache.
Bytes per parameter
These figures are approximate and count only the quantized tensors. The "calculated" figures come from each format's storage layout, with the arithmetic in the last column.
| Format | Bits per weight | Bytes per parameter | How it is stored |
|---|---|---|---|
| BF16 / FP16 | 16 | 2 | One 16-bit float per weight. This is the usual published checkpoint. Calculated. |
| FP8 (E4M3) | ≈ 8 | ≈ 1 | One 8-bit float per weight, plus a scale per tensor, channel or block. The scales add very little. Calculated. |
| GGUF Q8_0 | 8.5 | ≈ 1.06 | Each block of 32 weights shares one 16-bit scale: (32 × 8 + 16) ÷ 32 = 8.5. Calculated. |
| GGUF Q5_K_M | ≈ 5.7 | ≈ 0.71 | 5-bit K-quant blocks, with some tensors kept at 6-bit. Vendor figure: llama.cpp's quantize table gives 5.70 for Llama 3.1 8B. |
| GGUF Q4_K_M | ≈ 4.9 | ≈ 0.61 | 4-bit K-quant super-blocks (256 weights, 4.5 bits including their scales), with some tensors kept at 6-bit. Vendor figure: the same table gives 4.89. |
| NVFP4 | 4.5 | ≈ 0.56 | Each block of 16 FP4 values shares one FP8 scale: 4 + 8 ÷ 16 = 4.5. There is also one FP32 scale per tensor. Calculated. |
| INT4 AWQ / GPTQ, group size 128 | ≈ 4.16 | ≈ 0.52 | Each group of 128 weights shares a 16-bit scale and a 4-bit zero point: 4 + (16 + 4) ÷ 128 ≈ 4.16. Calculated. |
Real files come out larger than this table for two reasons:
- Many safetensors builds (FP8, AWQ, GPTQ, NVFP4) keep the embedding table and the output head at 16-bit, and some keep other layers at 16-bit too. GGUF quantizes those as well.
- The
_Mmixes keep chosen tensors at a higher precision.
Use the real file sizes. The Qwen3-8B check below shows how far they differ from the table.
MXFP4 is a close cousin of NVFP4 and is the format gpt-oss's expert weights ship in. It uses 32-value blocks with an 8-bit power-of-two scale: 4 + 8 ÷ 32 = 4.25 bits per weight (calculated from the format).
Quality: the method matters, not just the bits
The bit count limits how accurate a build can be, and the method decides how close it gets to that limit. Two "4-bit" builds of the same model can differ in:
- Number format. INT4 uses evenly spaced integers. NVFP4 uses a tiny float (E2M1), which has finer steps near zero, scaled per 16 values.
- Block or group size. Smaller blocks (16 in NVFP4, 32 in Q4_K's sub-blocks, 128 in typical AWQ/GPTQ) follow local value ranges more closely. The cost is a little more scale data per weight.
- How rounding is chosen. GPTQ uses calibration data to compensate for each rounding error in the weights that remain. AWQ rescales the channels that real activations rely on most before rounding. llama.cpp K-quants can use an importance matrix (imatrix) computed from sample text.
- What stays at higher precision. Each build decides about the embeddings, the output head and sensitive layers.
- Calibration data. A build calibrated only on English chat may lose more on code or other languages.
The community widely reports these patterns. They are tendencies, not rules:
- 8-bit (FP8, Q8_0) is usually very close to the 16-bit model.
- 5 and 6 bits lose little.
- At 4 bits, methods and models start to differ, most often in long reasoning, code, maths and exact recall.
- At the same bit width, smaller models tend to lose more than larger ones.
As one vendor figure: NVIDIA reports 1% or less accuracy loss going from FP8 to NVFP4 for DeepSeek-R1-0528 on the benchmarks it chose.
The test that matters is your own. Run the same set of real prompts against the 16-bit (or 8-bit) build and the quantized build, and compare. For GGUF, llama.cpp's llama-perplexity can also report the KL divergence from the full-precision model. That is a useful general signal, but it does not tell you how the build does on your task.
Speed: the runtime and the GPU decide
Fewer bytes only help if a kernel reads the packed weights directly. Which resource limits speed depends on the workload:
- Generating tokens for one or a few requests. Each new token reads every weight once, so memory bandwidth sets the pace. If the runtime has a kernel for the format, halving the bytes raises the ceiling to roughly twice the tokens per second (estimate).
- Reading a long prompt, or serving many requests at once. Arithmetic sets the pace. Weight-only formats such as AWQ and GPTQ INT4 (W4A16: 4-bit weights, 16-bit activations) unpack the weights to 16-bit before multiplying. That saves bandwidth but no arithmetic, and the unpacking takes time. FP8 on GPUs with FP8 tensor cores, and NVFP4 on Blackwell, also do the multiplication in low precision.
Where each format is accelerated (vendor figures from NVIDIA and vLLM documentation):
| Format | Native math on | On other GPUs |
|---|---|---|
| BF16 | NVIDIA Ampere and newer (RTX 30 series onward) | Older cards use FP16 instead |
| FP8 | Compute capability 8.9 and up: Ada (RTX 40, L4, L40S), Hopper (H100, H200), Blackwell | On Turing and Ampere, vLLM loads FP8 as weight-only (W8A16). You keep the memory saving, but the math runs in 16-bit. |
| NVFP4 | Blackwell: B200 and GB200, RTX 50 series, GB10 (DGX Spark) | Depends on the runtime. Some have weight-only fallback kernels and some do not load it at all. Check before you download. |
| INT4 AWQ / GPTQ | None: these formats are weight-only, so the math runs in 16-bit | vLLM runs them on Turing and newer, mostly through its Marlin kernels |
| GGUF (Q8_0 … Q4_K_M) | llama.cpp's own kernels for CUDA, Metal, Vulkan, ROCm and CPU | vLLM supports it experimentally |
Which runtime reads which format
| Build | File | Read by |
|---|---|---|
| BF16 / FP16 | safetensors | vLLM, SGLang, TensorRT-LLM, Hugging Face Transformers. llama.cpp can read it once converted to GGUF. |
| FP8 | safetensors plus scales, described in config.json | vLLM, SGLang, TensorRT-LLM |
| AWQ, GPTQ | safetensors, described in config.json | vLLM, SGLang, Transformers |
| NVFP4 | safetensors (an NVIDIA ModelOpt or llm-compressor export) | TensorRT-LLM, vLLM, SGLang, at full speed on Blackwell |
| GGUF Q8_0, Q5_K_M, Q4_K_M | one .gguf file holding the weights, tokenizer and metadata | llama.cpp and tools built on its GGML library (Ollama, LM Studio). vLLM: highly experimental, not optimized. |
These lists are not exhaustive. Check your runtime's support matrix for your GPU and version.
The picture
BF16 checkpoint (safetensors, 2 bytes per parameter)
|
+-- FP8 -----------+
+-- AWQ / GPTQ ----+--> safetensors + quantization_config
+-- NVFP4 ---------+ read by vLLM, SGLang, TensorRT-LLM
| FP8 math: Ada, Hopper, Blackwell
| FP4 math: Blackwell
|
+-- GGUF: Q8_0, Q5_K_M, Q4_K_M (one .gguf file)
read by llama.cpp, Ollama, LM Studio
CUDA, Metal, Vulkan, ROCm, CPU
NVFP4: 16 weights share one scale [4b][4b][4b] ... 16 values [FP8 scale] (16 × 4 + 8) ÷ 16 = 4.5 bits per weight weight ≈ value × block scale × tensor scale INT4 AWQ / GPTQ: 128 weights share one scale and zero point [4b][4b][4b] ... 128 values [16-bit scale] [4-bit zero] (128 × 4 + 16 + 4) ÷ 128 ≈ 4.16 bits per weight weight ≈ (q − zero) × scale
Worked example: one 27B model
Take a dense model with 27 billion parameters, the size class of Qwen3.8 27B. The arithmetic below rounds to 27 × 10⁹; take the exact count from the model's config or file listing. GB here means 10⁹ bytes, as Hugging Face lists file sizes. GPU memory is quoted in binary units, so a 12 GB card holds 12 × 1024³ ≈ 12.9 × 10⁹ bytes (calculated).
| Build | Bytes per parameter | Weights (calculated) |
|---|---|---|
| BF16 | 2 | 27 × 2 = 54 GB |
| GGUF Q8_0 | 1.0625 | 27 × 1.0625 ≈ 28.7 GB |
| FP8 | ≈ 1 | 27 × 1 = 27 GB |
| GGUF Q5_K_M | ≈ 0.71 | 27 × 0.71 ≈ 19.2 GB |
| GGUF Q4_K_M | ≈ 0.61 | 27 × 0.61 ≈ 16.5 GB |
| NVFP4 | 0.5625 | 27 × 0.5625 ≈ 15.2 GB |
| INT4 AWQ / GPTQ (group 128) | ≈ 0.52 | 27 × 0.52 ≈ 14.0 GB |
Real builds come out larger. NVIDIA's NVFP4 build of Qwen3.8 27B is 21.9 GB of files (vendor figure: the file listing on Hugging Face), not the 15.2 GB above: only its MLP layers are NVFP4, while its attention layers stay at FP8 and its vision encoder at BF16 (from the build's config). Add whatever the build keeps at higher precision. The embedding table and the output head each take vocabulary size × hidden size × 2 bytes, both from the config. Weights are also only part of the total: a model fits only if the weights, the KV cache, the runtime overhead and the working memory all fit.
What that means on three pieces of hardware:
- One RTX 3060 (12 GB). No build fits. Even the smallest 4-bit weights (≈ 14 GB, calculated above) are larger than the card before any KV cache. Offloading layers to the CPU lets the model run, but much more slowly. A 27B model is the wrong size for this card. Better choices are an 8B model at 8-bit (Qwen3-8B Q8_0: 8.7 GB file) or a 14B model at 4-bit (Qwen3-14B Q4_K_M: 9.0 GB file), both vendor figures from Qwen's repositories. They leave room for a modest context (estimate).
- Two RTX 3060s (2 × 12 GB).
- A runtime that supports it can split a 4-bit build (14–16.5 GB plus its 16-bit parts, calculated above) across both cards. llama.cpp splits by layers by default; vLLM uses tensor parallelism. That leaves a few GB per card for KV cache and overhead, so context is limited (estimate).
- Q5_K_M (≈ 19.2 GB) fits with less room for context (estimate). FP8 (27 GB) and BF16 do not fit.
- Two 12 GB cards are not one 24 GB card. Data crosses PCIe, and a single request does not get twice as fast. See One GPU vs multiple GPUs.
- These cards are Ampere, so NVFP4 gets no FP4 math and FP8 loads weight-only at best. The practical builds here are GGUF Q4_K_M and AWQ/GPTQ.
- DGX Spark (128 GB unified memory).
- Every build's weights fit, BF16 included, with room for KV cache. The operating system shares that memory.
- Here quantization matters for speed. Generating one token reads all the weights, so memory bandwidth ÷ weight bytes gives a ceiling for one stream. With NVIDIA's figure of 273 GB/s (vendor figure), the estimated ceilings are:
- BF16: 273 ÷ 54 ≈ 5 tokens/s
- FP8: 273 ÷ 27 ≈ 10 tokens/s
- NVFP4: 273 ÷ 15.2 ≈ 18 tokens/s
- GB10 is a Blackwell GPU, so FP8 and NVFP4 also get low-precision math.
- Serving many requests at once raises the total tokens per second, not the speed of a single stream.
Checking the arithmetic against real files: Qwen3-8B
Qwen publishes one model in several builds, so you can compare the arithmetic with real files. The parameter count, calculated from the published config (36 layers, hidden size 4,096, MLP size 12,288, 32 query heads and 8 KV heads of 128, vocabulary 151,936, separate output head):
- Attention per layer: 4,096 × (4,096 + 1,024 + 1,024) + 4,096 × 4,096 ≈ 41.9 M
- MLP per layer: 3 × 4,096 × 12,288 ≈ 151.0 M
- 36 layers × ≈ 192.9 M ≈ 6.95 B
- Embedding table and output head: 2 × 151,936 × 4,096 ≈ 1.24 B
- Total ≈ 8.19 B, the count the repository itself reports
| Build | Repository | Weight files (vendor figure) | Arithmetic (calculated) |
|---|---|---|---|
| BF16 | Qwen/Qwen3-8B | 16.4 GB | 8.19 × 2 ≈ 16.4 GB |
| FP8 | Qwen/Qwen3-8B-FP8 | 9.4 GB | 8.19 × 1 ≈ 8.2 GB, or 6.95 × 1 + 1.24 × 2 ≈ 9.4 GB with the embeddings and head at BF16 |
| AWQ INT4 | Qwen/Qwen3-8B-AWQ | 6.1 GB | 8.19 × 0.52 ≈ 4.3 GB, or 6.95 × 0.52 + 1.24 × 2 ≈ 6.1 GB |
| GGUF Q8_0 | Qwen/Qwen3-8B-GGUF | 8.7 GB | 8.19 × 1.0625 ≈ 8.7 GB |
| GGUF Q5_K_M | Qwen/Qwen3-8B-GGUF | 5.85 GB | 8.19 × 0.71 ≈ 5.8 GB |
| GGUF Q4_K_M | Qwen/Qwen3-8B-GGUF | 5.0 GB | 8.19 × 0.61 ≈ 5.0 GB |
The simple arithmetic matches the BF16 and GGUF files. It comes out 1.2 GB too low for FP8 and 1.8 GB too low for AWQ, because those two builds keep 1.24 B parameters at 16-bit. In a 27B model those parts are a smaller share of the total, but they are still there. Size a build from its files, not its name.
Try it
vLLM reads the quantization method from the checkpoint's config.json, so the command is the same as for the 16-bit model:
vllm serve Qwen/Qwen3-8B-AWQ --max-model-len 16384
Swap in Qwen/Qwen3-8B-FP8 or Qwen/Qwen3-8B to compare builds on your own prompts, on a GPU with room for them: the BF16 files alone are 16.4 GB (vendor figure). vLLM can also quantize a 16-bit checkpoint to FP8 while loading it. The option's scheme names have changed between versions, so check vllm serve --help.
llama.cpp loads a GGUF file. -ngl 99 asks for every layer on the GPU (recent builds choose automatically when you leave it out), and -c sets how much context the server allocates.
llama-server -m Qwen3-8B-Q4_K_M.gguf -c 16384 -ngl 99
To make a different level yourself, start from a 16-bit GGUF. llama.cpp's convert_hf_to_gguf.py creates one from a Hugging Face checkpoint. Then run:
llama-quantize Qwen3-8B-BF16.gguf Qwen3-8B-Q5_K_M.gguf Q5_K_M
The common mistake
The common mistake is treating "4-bit" as one thing: assuming every 4-bit build of a model has the same quality and runs faster everywhere.
- Quality differs by method. Q4_K_M, AWQ, GPTQ and NVFP4 round and group weights differently, and keep different layers at higher precision. Even two AWQ builds of the same model can differ because of their calibration data and group size.
- Speed needs a kernel.
- An NVFP4 build on an Ampere card gets no FP4 math and may not load at all.
- AWQ under heavy concurrency still multiplies in 16-bit.
- A GGUF file in vLLM runs on an experimental path.
- A 4-bit build that spills onto the CPU is far slower than a smaller model that fits on the GPU.
- Size comes from the files. A 4-bit build is not a quarter of the BF16 size, because embeddings, output heads and mixed tensors add to it.
Instead, pick a format your runtime and GPU accelerate, size it from its real files plus the KV cache, and compare it with the 8- or 16-bit build on your own prompts before committing to it.
On AxForge
- Find a model in the catalogue: language models, all models.
- Try GGUF and AWQ/GPTQ builds on 12 GB Ampere cards with the RTX 3060. You can use one card, or both with a model split across them.
- Try FP8 and NVFP4 builds on Blackwell with 128 GB of unified memory on the DGX Spark.
- Rent a machine: GPU rentals. To call a hosted model without choosing a build, see Quickstart and Models.
Next
- Context windows and KV cache: the other half of the memory bill.
- Model size, GPU memory and CPU offload: whether a build fits, and what happens when it does not.
- llama.cpp vs vLLM: which runtime suits your workload.