AI Engineering / Models & Memory

Model size, GPU memory and CPU offload

A model runs at GPU speed only when its weights, its KV cache, the runtime's overhead and its working memory all fit in GPU memory at once; when they do not, part of the model moves to system RAM, the CPU computes that part, and every token waits for it.

Will this model fit on my GPU?

Four things share GPU memory while a model serves requests:

PartWhat it isGrows with
WeightsThe model's parameters, loaded once at start.Parameter count × bytes per parameter of the build you load.
KV cacheThe keys and values that attention keeps for every token in context.Tokens in context × requests held at the same time.
Runtime overheadThe CUDA context, kernels and the runtime's own buffers.The runtime. Roughly fixed.
Working memoryActivations and scratch space while a batch is computed.Batch size, and how much of a prompt is processed at once.

Requirement = weights + KV cache + runtime overhead + working memory. If the sum is below the memory you have and leaves some headroom, the model fits. If it is above, the runtime has to give way somewhere: it refuses to start, shortens the context, or moves layers to the CPU.

Budget 1–1.5 GB for overhead and working memory on a single consumer GPU (estimate). The real figure depends on the runtime, the batch size, and whether a desktop session also uses the card.

Parameter count is not memory

"8B" counts parameters, not bytes. The number of bytes depends on the precision of the build. Llama 3.1 8B has 8.03 billion parameters (calculated from the model's published config):

PrecisionBytes per parameterWeightsSource
FP32432.1 GBcalculated: 8.03 × 10⁹ × 4
BF16 / FP16216.1 GBcalculated: 8.03 × 10⁹ × 2
8-bit (FP8, INT8)18.0 GB + scalescalculated: 8.03 × 10⁹ × 1
GGUF Q8_0≈ 1.06 (8.5 bits)8.54 GBpublished file size
4-bit (INT4, NVFP4)0.5 + scales4.0 GB + scalescalculated: 8.03 × 10⁹ × 0.5
GGUF Q4_K_M≈ 0.61 (4.9 bits)4.92 GBpublished file size
  • Real 4-bit files are bigger than parameters × 0.5. Each small block of weights carries a scale, and some tensors stay at higher precision. NVFP4, for example, stores one 8-bit scale per 16 four-bit values: 4 + 8 ÷ 16 = 4.5 bits per weight (calculated from the format).
  • Formats belong to runtimes. GGUF (Q4_K_M, Q8_0, …) is the llama.cpp world, including tools built on it such as Ollama. vLLM and similar servers load safetensors checkpoints in BF16, AWQ, GPTQ, FP8 or NVFP4 (vLLM's GGUF support is experimental).
  • Smaller is not automatically faster. A 4-bit build always saves memory and memory traffic. Whether it also computes faster depends on the runtime's kernels and the GPU. Blackwell GPUs such as GB10 accelerate NVFP4 in hardware. An Ampere card such as the RTX 3060 has no FP8 or FP4 tensor cores.
  • Quality varies by method. Two quantization methods at the same bit width can score differently on your task. Test the build you plan to run.
  • Mixture-of-experts models need every expert in memory but read only the active experts for each token. Size memory by the total parameter count, and expect single-request speed closer to what the active count suggests.

Why is my model using the CPU?

It is almost always one of two reasons:

  1. It does not fit, so the runtime offloaded part of it. Layers that do not fit in VRAM are kept in system RAM and computed by the CPU.
  2. The runtime cannot see the GPU at all. Common causes are a CPU-only build (llama.cpp compiled without CUDA, or a CPU-only PyTorch wheel), a container started without GPU access (start it with docker run --gpus all), or a driver the runtime does not support. In that case nvidia-smi lists no process for your runtime.

On Windows there is a third: by default the NVIDIA driver lets CUDA spill into system RAM when VRAM runs out (NVIDIA Control Panel, "CUDA - Sysmem Fallback Policy"), so an oversized model runs slowly instead of failing. "Prefer No Sysmem Fallback" turns that into an out-of-memory error.

What each runtime does when the model does not fit:

RuntimeWhen the model does not fit
llama.cpp (llama-server)Current builds fit the model automatically (--fit on, the default), changing only what you did not set. They shrink the context first, to no less than 4,096 tokens (vendor figure: the --fit-ctx default), and then move layers to system RAM. If you set -ngl, exactly that many layers go on the GPU and the rest run on the CPU.
OllamaSplits the model automatically. ollama ps shows the split.
Hugging Face Transformers with device_map="auto"Fills the GPU, puts the rest on the CPU and then on disk, and logs that parameters were offloaded.
vLLMDoes not offload unless you ask for it (--cpu-offload-gb). If the weights do not fit, loading fails with a CUDA out-of-memory error. If the KV cache is too small for --max-model-len, startup stops and reports the longest context that would fit.

Offload is slow because generating one token reads every weight once, so one request can produce at most memory bandwidth ÷ bytes read per token. Weights in system RAM can only be read at RAM speed:

  • RTX 3060 GDDR6: 360 GB/s (calculated: 15 Gbps × 192-bit bus ÷ 8).
  • Dual-channel DDR4-3200, a common desktop setup: 51.2 GB/s (calculated: 3,200 MT/s × 8 bytes × 2 channels).
  • PCIe 4.0 x16 between the two: 31.5 GB/s each way (calculated from the spec: 16 GT/s × 16 lanes × 128/130 ÷ 8). That is slower than reading the RAM directly, so streaming the offloaded weights to the GPU for every token would not help, and runtimes compute those layers on the CPU instead.

The GPU finishes its layers quickly and then waits for the CPU. That is why "my model is using the CPU" and "my GPU looks idle" are usually the same problem.

One token through Llama 3.1 8B on one 12 GB RTX 3060 (calculated upper bounds)

BF16 build, 16.1 GB: does not fit.
About 9 GB of layers on the GPU, about 7 GB in system RAM.

          +---------------------+      +---------------------+
token --> | GPU: first layers   | ---> | CPU: the rest       | --> next token
          | 9 GB / 360 GB/s     |      | 7 GB / 51.2 GB/s    |
          | ~25 ms              |      | ~137 ms             |
          +---------------------+      +---------------------+
          GPU busy ~25 of ~162 ms (15 %)   at most ~6 tokens/s

Q4_K_M build, 4.92 GB: fits. Every layer on the GPU.

          +---------------------+
token --> | GPU: all layers     | --> next token
          | 4.92 GB / 360 GB/s  |
          | ~14 ms              |
          +---------------------+
          GPU busy most of the token       at most ~73 tokens/s

These are upper bounds calculated from the bandwidth figures above, using a rounded, illustrative split. Real speeds are lower. They are also limits for a single request. A server that batches many requests reads each weight once for the whole batch, so total throughput across users is a different and larger number.

VRAM and unified memory

A discrete GPU such as the RTX 3060 has its own memory: 12 GB GDDR6 per card (vendor figure). System RAM is on the other side of PCIe. Here "fits" means fits in VRAM, and anything that does not fit is offloaded.

Unified memory gives the CPU and GPU one shared pool. On NVIDIA DGX Spark, built on the GB10 chip, that pool is 128 GB of LPDDR5x at 273 GB/s (vendor figures). There is no separate VRAM and no PCIe hop. That has three consequences:

  • The GPU can use most of the 128 GB, so models far too big for a consumer card fit.
  • Offloading to the CPU frees nothing, because it is the same memory. A build that does not fit the pool does not fit the machine.
  • The pool also holds the operating system and every other process, so leave headroom. When the pool runs out, the machine starts swapping and slows down as a whole, instead of one program failing cleanly. On GB10, nvidia-smi shows "Memory-Usage: Not Supported" (NVIDIA documents this for integrated GPUs), so read the pool with free -h.
One RTX 3060Two RTX 3060DGX Spark (GB10)
Memory (vendor figure)12 GB VRAM2 × 12 GB VRAM, two separate pools128 GB unified, shared with the CPU
Bandwidth360 GB/s (calculated)360 GB/s per card (calculated)273 GB/s (vendor figure)
Llama 3.1 8B, BF16 (16.1 GB, calculated)NoSplit across both cardsYes
Llama 3.1 8B, Q4_K_M (4.92 GB, published file size)YesYes, or one copy per cardYes
Llama 3.1 70B at 4-bit (35.3 GB before scales, calculated: 70.6 × 10⁹ × 0.5)NoNoYes

Two 12 GB cards are not one 24 GB card. A runtime that supports it can split one model across both cards, but the data passing between them crosses PCIe, and a single request does not get twice as fast. The multiple-GPU guide explains why.

Memory decides what fits, and bandwidth sets the ceiling on tokens per second for one request. The Spark holds the BF16 build easily, but reading 16.1 GB per token at 273 GB/s limits one request to about 17 tokens/s. The 4.92 GB 4-bit build on one RTX 3060 is limited to about 73 tokens/s. Both are calculated upper bounds from the bandwidth figures above.

Worked example: Llama 3.1 8B on one 12 GB RTX 3060

The model's inputs, from its published config: 32 layers, 8 KV heads, head dimension 128, native context 131,072 tokens. The card has 12 GB, which nvidia-smi shows as 12,288 MiB (vendor figure). In the decimal units that file sizes use, that is 12.9 GB (calculated: 12,288 × 1,048,576 bytes).

Units: Hugging Face lists file sizes in decimal GB (10⁹ bytes), and nvidia-smi reports MiB. 1 GiB = 1.074 GB.

1. Weights

At BF16 the weights are 8.03 × 10⁹ × 2 bytes = 16.1 GB (calculated). That is already larger than the card before any context, so this build cannot fit. The Q4_K_M build is 4.92 GB (published file size).

2. KV cache

The cache per token at FP16 (calculated from the model's published config):

2 (K and V) × 32 layers × 8 KV heads × 128 dims × 2 bytes = 131,072 bytes = 128 KiB per token

Context, one requestKV cacheSource
8,192 tokens1.07 GB (1 GiB)calculated: 131,072 bytes × 8,192
32,768 tokens4.29 GB (4 GiB)calculated: 131,072 bytes × 32,768
131,072 tokens (native maximum)17.2 GB (16 GiB)calculated: 131,072 bytes × 131,072

The cache is per request held at once. Two conversations at 32,768 tokens each need 8.6 GB (calculated).

3. Add it up

BuildContextWeightsKV cacheOverhead + working (estimate)TotalOn 12.9 GB
BF168,19216.11.07~1.5~18.7 GBNo, it offloads
Q8_08,1928.541.07~1.5~11.1 GBYes, ~1.8 GB spare
Q4_K_M8,1924.921.07~1.5~7.5 GBYes, ~5.4 GB spare
Q4_K_M32,7684.924.29~1.5~10.7 GBYes, ~2.2 GB spare
Q4_K_M131,0724.9217.2~1.5~23.6 GBNo

Sources: the BF16 weights and every KV figure are calculated from the published config. The GGUF weights are the published file sizes. Overhead is an estimate.

  • BF16 cannot run on this card without offload, and with offload it runs at roughly the speed in the diagram above.
  • Q4_K_M fits with room for a 32K context for one request.
  • 131,072 tokens is what the model can handle, not what you should allocate. Both llama-server (-c 0, the default) and vLLM (--max-model-len unset) start from the model's maximum, so set the context you need.
  • Two requests at 32K would need 4.92 + 8.6 + ~1.5 ≈ 15 GB, which no longer fits on one card. Hold fewer tokens, use an 8-bit KV cache, or add memory. In llama-server, -c is the cache for all requests together, so two full 32K conversations need -c 65536.

To run the build that fits, with the context you chose and every layer on the GPU:

llama-server -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99

-c sets how much context is allocated. -ngl 99 asks for every layer on the GPU (99 is more layers than the model has; current builds also accept -ngl all). Once -ngl and -c are set, llama.cpp no longer adjusts them on its own. If the model and cache do not fit, loading fails with an out-of-memory error rather than quietly running part of the model on the CPU (unless the Windows driver fallback above is on). In vLLM the same choice is a 4-bit checkpoint plus --max-model-len.

Architecture matters as much as size. Qwen2.5-7B has 28 layers and 4 KV heads (published config). Its cache is 2 × 28 × 4 × 128 × 2 = 57,344 bytes per token, or 1.88 GB at 32,768 tokens (calculated), less than half of Llama 3.1 8B's. Models with sliding-window or hybrid attention keep a full-length cache in only some layers, and MLA models (DeepSeek) store a compressed one, so this formula overstates their cache. Check the architecture before you apply it. Context and the KV cache covers those models.

How to see offload

Watch the GPU while tokens are being generated. If utilization stays low while top (Task Manager on Windows) shows the CPU cores busy, the GPU is waiting on the CPU.

nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv -l 1

A full card on its own proves nothing: vLLM reserves most of the GPU's memory by default (--gpu-memory-utilization: 0.9 for a long time, 0.92 in current releases) and fills what the weights leave with KV cache, and llama-server allocates the whole -c cache at start.

llama.cpp: read the load log. Llama 3.1 8B has 33 layers that can go to the GPU: 32 blocks plus the output layer. offloaded 33/33 layers to GPU means the whole model is on the GPU. Any lower number means the rest runs on the CPU. When automatic fitting is on, the log also says what it changed, such as a shortened context.

Ollama: the PROCESSOR column shows 100% GPU when everything is on the card, and a split such as 48%/52% CPU/GPU when it is not (example values from Ollama's docs).

ollama ps

Transformers: print which device each layer landed on.

from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    device_map="auto",  # fills the GPU, then CPU RAM, then disk
    dtype="auto",       # the checkpoint's precision (BF16), not FP32; transformers 4.56+
)
print(model.hf_device_map)  # any "cpu" or "disk" value is an offloaded layer

vLLM does not offload by default. Its startup log reports how much memory the weights took and how many tokens of KV cache fit.

The common mistake

The model is slow and nvidia-smi shows the GPU mostly idle, so the conclusion is "this GPU is too small, rent a bigger machine". Usually the build is too big for the card, not the card too small for the job. Work through these steps in order:

  1. Pick a build that fits. That means a Q4_K_M or Q8_0 GGUF for llama.cpp, an AWQ, GPTQ or FP8 checkpoint for vLLM, or NVFP4 on Blackwell.
  2. Allocate the context you use, not the model's maximum (-c, --max-model-len).
  3. Shrink the KV cache if long context is what you need. An 8-bit cache roughly halves it: -fa on -ctk q8_0 -ctv q8_0 in llama.cpp, or --kv-cache-dtype fp8 in vLLM where the GPU and attention backend support it. Check the quality on your task.
  4. Only then add memory, when the build you need still does not fit. Typical cases are BF16 for quality, many long requests at once, or a 70B model. Two 12 GB cards give you a second 12 GB pool, not a faster 24 GB card. 128 GB of unified memory fits far larger models, at 273 GB/s (vendor figure).

On AxForge

Next

Ask on the forum Markdown For AI Updated 2026-10-01

Anything unclear on this page?

Ask on the forum — the answer helps the next person too.

Ask about this page