# llama.cpp vs vLLM


# llama.cpp vs vLLM

llama.cpp runs a model on whatever hardware you have for one or a few conversations at a time; vLLM serves many concurrent requests from GPUs that hold the whole model — so pick by how many requests will be in flight, not by a single-stream benchmark.

## Pick by workload

    | Workload | Start with | Why |  |

      | One developer chatting, trying models | llama.cpp | One program, one GGUF file, quick to start and to swap models |  |

      | A model larger than your GPU memory | llama.cpp | Layers that do not fit run on the CPU — much slower, but it runs |  |

      | Laptop, Mac or CPU-only machine | llama.cpp | CPU, Apple Metal, CUDA, ROCm and Vulkan backends |  |

      | A few users, or one agent, on one GPU | Either | llama.cpp with a few slots (`-np`) is often enough — measure at your concurrency |  |

      | Agent harnesses firing parallel and repeated requests | vLLM | KV cache handed out in blocks as tokens arrive; identical prompt starts shared across running requests |  |

      | An API or team server with steady concurrent traffic | vLLM | Total throughput and latency under load |  |

## What each runtime is built for

### llama.cpp

  - A C/C++ inference engine. `llama-server` is its OpenAI-compatible HTTP server.

  - Models come as **GGUF** files — weights, tokenizer and chat template in one file, usually quantized (Q4_K_M, Q5_K_M, Q8_0; see quantization).

  - Runs on CPU, Apple Metal, NVIDIA CUDA, AMD ROCm and Vulkan.

  - **Layer offload:** `-ngl` sets how many layers live on the GPU; the rest run on the CPU from system RAM. A model bigger than VRAM still runs, much more slowly (model fit and CPU offload).

  - **Slots:** a fixed number of parallel sequences (`-np`), with continuous batching across the busy slots on by default. The context set with `-c` is the total for all slots: split evenly between them, or one shared pool with `--kv-unified` (current builds turn that on only when the slot count is left on auto). The default slot count has changed between versions, so set it explicitly.

  - **Prompt cache:** a slot reuses the part of a new prompt that matches what it already holds, so a growing conversation is not recomputed from scratch each turn. Recent builds also keep idle slots' prompts in host RAM (`--cache-ram`) to restore later. A slot cannot reuse what a busy slot holds.

### vLLM

  - A Python serving engine built on PyTorch. `vllm serve` starts an OpenAI-compatible server. Main targets: Linux with NVIDIA or AMD GPUs.

  - Models come as Hugging Face checkpoints: BF16/FP16, or quantized as FP8, AWQ, GPTQ or NVFP4 (NVFP4 is hardware-accelerated on Blackwell GPUs such as the GB10). GGUF loading exists but is documented as highly experimental.

  - **Paged KV cache** (PagedAttention): at startup vLLM reserves a share of GPU memory (`--gpu-memory-utilization`, about 0.9 by default per the docs). What weights and working memory leave becomes a pool of fixed-size KV blocks, handed to sequences as their tokens arrive and returned when they finish.

  - **Continuous batching:** every step the scheduler decides which sequences run; finished ones leave and waiting ones join without waiting for the batch to drain. `--max-num-seqs` caps how many run at once.

  - **Automatic prefix caching:** KV blocks for an identical prompt start — a system prompt, a tool list, a shared document — are computed once and reused by every request that begins the same way, including requests still running. On by default in current versions for models that support it.

  - Built for weights that sit in GPU memory. `--cpu-offload-gb` exists, but the offloaded weights then cross PCIe on every step. Several GPUs work through tensor parallelism (`--tensor-parallel-size`), which adds traffic between the cards (one GPU vs multiple GPUs).

## Why the number of requests changes the comparison

Generating one token (decode) reads every weight of a dense model once. With a single stream the GPU spends most of that time waiting on memory, not computing, so memory bandwidth sets the ceiling:

```
tokens/s for one stream  ≤  memory bandwidth ÷ bytes read per token

RTX 3060 12 GB with Qwen2.5-7B-Instruct Q4_K_M:
  360 GB/s ÷ 4.68 GB  ≈  77 tokens/s   (a rough ceiling)
```

360 GB/s: calculated from the 12 GB card's memory spec — 15 Gbps GDDR6 on a 192-bit bus, 15 × 192 ÷ 8 = 360 (NVIDIA lists the 192-bit bus; board makers list the 15 Gbps). 4.68 GB: the published size of the Q4_K_M GGUF file on Hugging Face, used as the bytes read per token. 77 tokens/s: calculated ceiling, not a speed you will see — the KV cache is read too, and no kernel reaches peak bandwidth.

Batching changes the arithmetic. One pass over those 4.68 GB can produce one token for each of eight sequences. Total tokens per second climbs; each stream runs a little slower than it would alone, more so as the batch and its KV reads grow. Keeping enough sequences in every step is what a serving engine is for — and a one-stream test cannot see it.

## How requests reach the model

```
ONE DEVELOPER, ONE CONVERSATION

  chat window or editor
          │  one request at a time
          ▼
  llama-server
    slot 0 ── KV cache for this conversation
          │
          ▼
  model weights (GGUF)
    all layers on the GPU, or some on the CPU (-ngl)
          │
          ▼
  tokens stream back, one per step
          │
          ▼
  you read, think, type … the GPU waits
```

```
MANY REQUESTS IN FLIGHT

  main agent turn ─┐
  sub-agent 1 ─────┤
  sub-agent 2 ─────┼──► request queue
  colleague ───────┤          │
  CI job ──────────┘          ▼
  ┌─────────────────────────────────────────────┐
  │ scheduler: which sequences run this step    │
  │   prefix cache: identical prompt starts     │
  │     computed once, shared by reference      │
  │   paged KV cache: blocks handed out as      │
  │     tokens arrive, returned when done       │
  └──────────────────────┬──────────────────────┘
                         ▼
  one batched pass over the model weights
  = one new token for every running sequence;
  finished sequences leave, waiting ones join
```

llama-server can sit in the second picture too: with `-np 4` it batches four slots. What differs is how far it goes — a fixed slot count, a context budget split per slot (or pooled with `--kv-unified`), and prefix reuse only from idle slots, against a scheduler that admits sequences while KV blocks last and shares identical prefixes across running requests.

## Workload fit: no runtime fills a GPU by itself

vLLM does not make a GPU busy; requests do. How close either runtime gets to the hardware depends on how many sequences are in flight and how long they are.

```
█ prefill (reading the prompt)   ▒ decode (writing tokens)   · idle
illustrative, not to scale

one developer chatting
GPU  ██▒▒▒▒▒▒··························██▒▒▒▒▒▒▒▒························
     ask, answer   you read and type   ask, answer   you read and type

agent loop with four sub-agents
GPU  ██▒▒█▒▒▒██▒▒▒█▒▒██▒▒▒▒█▒▒▒██▒▒█▒▒▒▒██▒▒▒█▒▒██▒▒▒▒█▒▒▒██▒▒▒█▒▒▒▒██
     tool results arrive and new turns start while other turns decode
```

  - **In chat, the person is the gap.** While you read and type, the GPU idles, whatever runtime drives it.

  - **nvidia-smi's "GPU-Util" is not "fully used".** It is the share of time any kernel was running. One decoding stream can show a high number while the compute units mostly wait on memory.

  - **"Memory full" on an idle vLLM server** is its reserved KV pool, not load.

  - **Concurrency has two caps:** the slot or sequence limit (`-np`, `--max-num-seqs`) and KV memory. If 4 GiB is left for KV after weights and working memory (estimate for a 7B 4-bit model on a 12 GB card), Qwen2.5-7B at 56 KiB per token (calculated in the worked example) holds 4 GiB ÷ 56 KiB ≈ 74,900 tokens in total (calculated): four sequences of 16K, or eighteen of 4K. Past that, requests wait; vLLM can also pause a running sequence and recompute it later.

  - **Long prompts compete with decoding.** Prefill is compute-heavy. Both runtimes feed a long prompt in pieces (vLLM calls it chunked prefill) so the other streams keep producing tokens, but those streams slow down while it runs.

## Start each one

### llama.cpp

A 7B model in 4-bit on one GPU. Download the GGUF file (`Qwen2.5-7B-Instruct-Q4_K_M.gguf`, in the `bartowski/Qwen2.5-7B-Instruct-GGUF` repository on Hugging Face) first, or let `-hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M` fetch it in place of `-m`.

```
llama-server -m Qwen2.5-7B-Instruct-Q4_K_M.gguf -c 16384 -ngl 99
```

  - `-m` — the GGUF file.

  - `-c 16384` — total context for all slots. Set it: left unset, current builds start from the model's trained maximum and shrink it to fit memory. The trained maximum is what the model can handle, not what you need to reserve.

  - `-ngl 99` — every layer on the GPU (this model has 28; 99 is the usual way to say "all"). Current builds default to auto, which places as many as fit. Lower it and the remaining layers run on the CPU.

  - The server answers on `http://localhost:8080/v1`.

For several conversations at once, set the slot count. With `-np 4` the 32,768 tokens are split evenly: 8,192 per request (calculated: 32,768 ÷ 4). Add `--kv-unified` to let the four draw from one pool instead, so one long request can use more while the others are short.

```
llama-server -m Qwen2.5-7B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99 -np 4
```

### vLLM

The same model as a Hugging Face checkpoint, installed with `pip install vllm` on Linux. The BF16 weights alone are about 15.2 GB (7.61B parameters × 2 bytes, calculated) — more than a 12 GB card. The AWQ 4-bit build (5.57 GB of weight files, published size) fits:

```
vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ --max-model-len 16384
```

  - `--max-model-len 16384` — the longest prompt plus answer per request. Left unset, vLLM takes the model's maximum (32,768 in this model's config) and stops at startup with an error, naming the length that would fit, if the KV pool cannot hold one sequence that long. Current versions also accept `--max-model-len auto`, which picks the largest length that fits.

  - `--gpu-memory-utilization` and `--max-num-seqs` are the knobs for the KV pool and for concurrency.

  - The server answers on `http://localhost:8000/v1`.

Both expose `/v1/chat/completions`, so any OpenAI client works against either by changing its base URL (OpenAI SDK).

## Worked example: one developer vs an agent loop

Same model, Qwen2.5-7B-Instruct, on one GPU. Its KV cache per token, calculated from the model's published config (28 layers, 4 KV heads, head_dim 128) at FP16:

```
2 (K and V) × 28 layers × 4 KV heads × 128 head_dim × 2 bytes
  = 57,344 bytes = 56 KiB per token
```

### A: one developer chatting

  - One request in flight: a 2,000-token prompt, a 400-token answer, then a minute of reading (estimates).

  - Memory: 4.68 GB of weights + 16,384 × 56 KiB = 896 MiB of KV (calculated) + a few hundred MB of runtime buffers (estimate). Fits a 12 GB card with room to spare.

  - What you feel: time to first token and the speed of one stream — the bandwidth-bound regime where both runtimes sit under the same ceiling.

  - **Fit: llama.cpp.** One file, one command, swap models by restarting with another file. vLLM's scheduler and prefix sharing would have nothing to work on.

### B: an agent loop with tool calls

Assumptions (estimates): system prompt plus tool definitions = 8,000 tokens; each turn the model emits a tool call and the harness appends the result, about 1,000 new tokens per turn; 20 turns.

  - Prompt tokens sent over the run: 20 × 8,000 + 1,000 × (0 + 1 + … + 19) = 160,000 + 190,000 = **350,000**.

  - With the prefix reused, only new tokens need prefill: 8,000 + 19 × 1,000 = **27,000** — about 13× less prefill work (calculated from the assumptions). llama.cpp's slot cache gets this for a single loop too, as long as the loop stays on its slot.

  - Now the harness fans out four sub-agents (searching different directories) that start from the same 8,000-token prefix and add 3,000 tokens each, while the main loop is at 20,000 (calculated):

      - KV with no sharing: 20,000 + 4 × 11,000 = 64,000 tokens × 56 KiB ≈ 3.4 GiB

      - KV with the prefix stored once: 20,000 + 4 × 3,000 = 32,000 tokens × 56 KiB ≈ 1.7 GiB

  - Five sequences decode in the same steps; a colleague's agent joins mid-run without waiting for a free slot.

  - **Fit: vLLM.** Shared prefix blocks, KV that grows with each sequence, a scheduler that admits new requests every step. With llama.cpp you would set `-np 5` or more; without `--kv-unified` each slot gets a fifth of `-c`, so `-c` must cover five times the longest conversation. Either way each busy slot computes and stores its own copy of the 8,000-token prefix — the 3.4 GiB case, not the 1.7.

## The common mistake: one stream, one verdict

One prompt goes to each runtime, someone reads tokens per second and declares a winner. That test answers a narrow question:

  - It measures the bandwidth-bound regime, where runtimes are closest. Single-stream comparisons go both ways depending on model, format and GPU, and the gap in vLLM's favour tends to open as concurrency rises (community measurements).

  - It often compares formats, not runtimes: a 4.68 GB Q4_K_M file reads about 3.3× fewer bytes per token than the 15.2 GB BF16 checkpoint (calculated: 15.2 ÷ 4.68).

  - It never exercises the scheduler, the paged KV cache or the prefix cache — the parts that decide how an agent harness or a team server behaves.

Measure the workload you will run:

    | Metric | What it shows | Matters most for |  |

      | Time to first token | Queueing plus prefill | Chat feel; agents with long prompts |  |

      | Tokens/s per stream | Decode speed one user sees | Interactive chat |  |

      | Total tokens/s | Work done across all requests | Agents, teams, APIs |  |

      | p95 latency at your concurrency | How slow the slow requests get | Anything shared |  |

  - Same model at the same precision — or write the difference down.

  - Same context limit on both.

  - The concurrency you expect (1, 4, 16 requests in flight), real prompt lengths, and repeated prefixes if agents are the load.

  - A load generator that holds N requests open: vLLM ships `vllm bench serve` (with `--max-concurrency`), and it can also target other OpenAI-compatible servers.

## On AxForge

  - The API at `https://api.axforge.ai/v1` is OpenAI-compatible, like llama-server and vLLM, so the same client code works against all three: Quickstart, OpenAI SDK.

  - Point a coding agent at an OpenAI-compatible endpoint: aider, Cline, Continue, Codex and Claude Code.

  - Run vLLM or llama.cpp yourself on a dedicated machine with SSH access: GPU rentals — RTX 3060 (12 GB per card, one card or both) and DGX Spark (128 GB unified memory).

  - Models: language models in the catalogue and the models on the API.

## Next

  - One GPU vs multiple GPUs — when a second card adds capacity and when it adds throughput.

  - Context windows and the KV cache — the per-token arithmetic behind every concurrency limit on this page.

  - How coding agents work — why an agent loop sends so many, so similar requests.



Source: https://dev.axforge.ai/ai-engineering/llama-cpp-vs-vllm/
