# One GPU vs multiple GPUs


# One GPU vs multiple GPUs

A second GPU gives a split model more room or runs a second independent stream, but two cards stay two separate memories joined by a link, and one request does not get twice as fast.

## Two cards are two memories

Each GPU has its own memory. A machine with two RTX 3060s has 12 GB + 12 GB (vendor figure; `nvidia-smi` shows 12,288 MiB per card). CUDA programs can use about 11.75 GiB of each card (community measurement: the total PyTorch reports on a 12 GB RTX 3060). That is not one 24 GB card. Software sees two separate devices.

  - **One model across both cards needs a runtime that can split it.** llama.cpp and vLLM both can. Anything the two halves exchange travels over PCIe, because the RTX 3060 has no NVLink connector.

  - **Each card has to fit on its own.** Every card pays its own runtime overhead and holds its own share of the weights and the KV cache. A model loads only if its share fits on the fullest card.

  - **Not every tool splits a model.** Many image and video pipelines run each model on one GPU by default. With those tools, the second card only helps by running a second job.

## Three ways to use two GPUs

    | Way | On each card | Crosses the link | What you gain | Flags |  |

    | **Layer split**
pipeline parallelism | A block of consecutive layers and their KV cache | The hidden state at the boundary, once per token | Room for a bigger model. One request runs about as fast as on one larger card with the same memory bandwidth, because the cards take turns. | llama.cpp `--split-mode layer` (the default)
vLLM `--pipeline-parallel-size 2` |  |

    | **Tensor split**
tensor parallelism | A slice of every layer's weights and of the KV cache | Partial results, combined by an all-reduce twice per layer | Room for a bigger model. One request can get faster, but by less than 2×. How much depends on the link. | llama.cpp `--split-mode tensor` (experimental)
vLLM `--tensor-parallel-size 2` |  |

    | **Replicas**
data parallelism | A whole model | Nothing | Two independent streams, each at one card's speed. The model has to fit on one card. | One server per card (`CUDA_VISIBLE_DEVICES`)
vLLM `--data-parallel-size 2` |  |

llama.cpp also has an older `--split-mode row`. Its own documentation marks it deprecated.

## How the work moves

```
LAYER SPLIT  (pipeline parallelism)
one model cut between layers; the cards take turns

  prompt --> [ GPU 0 | layers 1-32  + their KV ]
                   |  hidden state, ~10 KiB per token
                   v
             [ GPU 1 | layers 33-64 + their KV ] --> token

TENSOR SPLIT  (tensor parallelism)
every layer cut in two; the cards work at once, then sync

            +-> [ GPU 0 | half of every layer ] -+
  prompt ---+          ^  all-reduce             +--> token
            +-> [ GPU 1 | half of every layer ] -+
                       v  2x per layer = 128 syncs per token

REPLICAS  (two servers, or data parallelism)
a whole model on each card; nothing crosses between them

  request A --> [ GPU 0 | whole 8B model ] --> answer A
  request B --> [ GPU 1 | whole 8B model ] --> answer B
```

The sizes are for Qwen2.5-32B (64 layers, hidden size 5,120), calculated from the model's published config at 2 bytes per value.

## The link: PCIe vs NVLink

How much data has to cross between the cards depends on the split. All figures below are calculated from Qwen2.5-32B's published config, at 2 bytes per value:

  - **Layer split, generating:** one hidden state crosses per token: 5,120 × 2 B = 10 KiB. Any link carries that easily.

  - **Tensor split, generating:** 2 all-reduces per layer × 64 layers = 128 synchronisations per token, each about 10 KiB. The volume is small. The cost is the 128 waits, so link latency and the runtime's all-reduce code decide how much speed you keep.

  - **Tensor split, reading a 4,096-token prompt:** each all-reduce carries 4,096 × 5,120 × 2 B = 40 MiB. Over 128 all-reduces, each card sends about 5 GiB, because a two-card ring all-reduce sends roughly the whole message per card.

    | Link | Bandwidth, one direction | Moving those 5 GiB (calculated, ideal) |  |

    | PCIe 4.0 ×16 | 31.5 GB/s: 16 GT/s × 16 lanes × 128/130 ÷ 8 bits (calculated from the PCIe 4.0 spec) | ≈ 0.17 s |  |

    | PCIe 4.0 ×4 | 7.9 GB/s (same arithmetic, 4 lanes) | ≈ 0.68 s |  |

    | NVLink on an H100 SXM | 450 GB/s (vendor figure: 900 GB/s in both directions together) | ≈ 0.012 s |  |

vLLM's documentation recommends pipeline parallelism over tensor parallelism on GPUs without NVLink, for higher throughput and lower communication overhead. llama.cpp defaults to the layer split, and its documentation says the tensor mode is much more bottlenecked by the GPU interconnect. Over PCIe, a tensor split can still speed up a single stream. Measure it on your own workload before you commit to it.

Many desktop boards run a second slot at ×8 or ×4. This command shows each card's current link and memory. Run it while the cards are busy, because an idle card drops to a lower PCIe generation to save power:

```
nvidia-smi --query-gpu=index,name,memory.used,memory.total,pcie.link.gen.current,pcie.link.width.current --format=csv
```

## Capacity or throughput

These two measures are often confused:

  - **Single-stream speed** is the tokens per second that one request sees.

  - **Aggregate throughput** is the tokens per second across all requests running at once.

A second card can raise one of them, both, or neither. It depends on whether your model fits on one card and on how many requests arrive at once:

    | Your situation | Use the second card for | What changes |  |

    | The model does not fit on one card | A split (layer or tensor) | The model loads. Single-stream speed stays near what one larger card with the same bandwidth would give (layer split) or rises somewhat (tensor split). |  |

    | The model fits, and many requests arrive at once | Replicas | Aggregate throughput rises, up to about twice (estimate: two independent copies that share only the host). Each single request runs at the same speed as before. |  |

    | The model fits, and requests come one at a time (one developer chatting) | A second job, or a tensor split | With replicas, nothing changes: each request runs on one card while the other idles. A tensor split can make each request somewhat faster; over PCIe, measure it first. |  |

    | The model fits, but long contexts or many users run out of KV memory | A split | Each card holds only half the weights, which leaves more room for the KV cache. That means more context or more concurrent sequences. |  |

**Why one request does not get 2× faster.** Generating a token reads every weight once, so single-stream speed is at most roughly memory bandwidth ÷ bytes of weights.

  - **Layer split:** the two halves are read one after the other, so the total time per token stays the same.

  - **Tensor split:** the halves are read at the same time, but then the cards wait for each other 128 times per token.

  - **Replicas:** each request runs on a single card.

One card can also serve several requests at once. llama-server runs parallel slots (`-np`), and vLLM batches requests continuously, so one read of the weights serves every sequence in the batch. A second card is what you add when one card's memory or bandwidth becomes the limit. It is not the only way to serve two users.

More capacity can also come from one device with more memory. A DGX Spark has 128 GB of unified memory on a single GB10 chip (vendor figure), so a 20 GB model needs no split and no link. Which setup is faster depends on two things:

  - memory bandwidth, which sets single-stream generation speed: 273 GB/s on a DGX Spark, 360 GB/s per RTX 3060 (vendor figure; calculated);

  - compute, which matters for long prompts and many simultaneous requests.

## Worked example: one ~20 GB model split, or two 8B models side by side

### A. Qwen2.5-32B-Instruct at Q4_K_M, split across both cards

The GGUF file is 19,851,336,576 bytes, which is 18.5 GiB (published file size on Hugging Face for `bartowski/Qwen2.5-32B-Instruct-GGUF`). That is too big for one card. A layer split puts about half of the model's 64 layers on each card. Per card:

    | Per card | Size | Basis |  |

    | Memory CUDA can use | ≈ 11.75 GiB | community measurement |  |

    | Weights: half the file | ≈ 9.25 GiB | calculated |  |

    | Runtime overhead: CUDA context, compute buffers | ≈ 0.75 GiB | estimate |  |

    | Left for the KV cache: 11.75 − 9.25 − 0.75 | ≈ 1.75 GiB | calculated from the rows above |  |

    | KV per token on this card: 2 (K, V) × 32 layers × 8 KV heads × 128 head_dim × 2 B | 128 KiB | calculated from the model's published config |  |

    | Context that fits: 1.75 GiB ÷ 128 KiB | ≈ 14,000 tokens | estimate (moves with the overhead) |  |

With `-c 8192`, the KV cache takes 1 GiB per card and the card uses about 11 GiB (calculated, with the overhead estimate), so the model loads with room to spare. The model's native window is 32,768 tokens (`max_position_embeddings`, from its published config). That would need 4 GiB of KV cache per card at 2 bytes per value (f16, llama.cpp's default), which does not fit. Real splits are not exactly even, because 64 layers plus the output layer do not divide into two equal halves (estimate: a few hundred MiB apart). If one card runs out first, shift layers with `--tensor-split`, for example `52,48`.

```
llama-server -m Qwen2.5-32B-Instruct-Q4_K_M.gguf -c 8192 -ngl 99 --split-mode layer --tensor-split 1,1
```

To try llama.cpp's tensor split, use `--split-mode tensor -fa on` instead of the last two flags. It is experimental, it needs flash attention, and it needs an unquantized KV cache (f16, the default).

In vLLM, the official 4-bit AWQ build is 19,328,993,904 bytes, which is 18.0 GiB (published file size on Hugging Face), or 9.0 GiB per card (calculated).

  - `--gpu-memory-utilization 0.95` lets vLLM use 95% of what CUDA reports: 0.95 × 11.75 ≈ 11.2 GiB per card. That leaves ≈ 2.2 GiB for activations, CUDA graphs and the KV cache (calculated). This is tight. If vLLM stops with a KV-cache error, lower `--max-model-len`. Its startup log reports how many tokens of KV cache it got.

  - Tensor parallelism needs the attention-head count to be divisible by the GPU count. Qwen2.5-32B has 40 heads (published config), so 2 GPUs work.

```
vllm serve Qwen/Qwen2.5-32B-Instruct-AWQ --tensor-parallel-size 2 --max-model-len 8192 --gpu-memory-utilization 0.95
```

For a pipeline (layer) split in vLLM, replace `--tensor-parallel-size 2` with `--pipeline-parallel-size 2`.

**Speed ceiling (estimate).** The RTX 3060 has 360 GB/s of memory bandwidth (calculated: 15 Gbps × 192-bit bus ÷ 8), and every weight is read once per token.

  - **Layer split:** 19.85 GB ÷ 360 GB/s ≈ 55 ms per token, so at most ≈ 18 tokens/s. A single larger card with the same bandwidth would hit the same ceiling.

  - **Tensor split:** each card reads its half at the same time, ≈ 28 ms per token, so at most ≈ 36 tokens/s before the 128 waits per token.

  - **For comparison, on one DGX Spark:** 19.85 GB ÷ 273 GB/s ≈ 73 ms per token, so at most ≈ 14 tokens/s, with no split and far more room for context.

Real runs land below these ceilings.

### B. Two 8B models, one per card

Llama 3.1 8B Instruct at Q8_0 is 8,540,775,840 bytes, which is 7.95 GiB (published file size on Hugging Face for `bartowski/Meta-Llama-3.1-8B-Instruct-GGUF`). It fits on one card:

    | Per card | Size | Basis |  |

    | Weights | ≈ 7.95 GiB | calculated |  |

    | KV per token: 2 × 32 layers × 8 KV heads × 128 head_dim × 2 B | 128 KiB | calculated from the model's published config |  |

    | KV for `-c 16384`: 16,384 × 128 KiB | 2 GiB | calculated |  |

    | Runtime overhead | ≈ 0.75 GiB | estimate |  |

    | Total, out of ≈ 11.75 GiB | ≈ 10.7 GiB | estimate |  |

```
# one per terminal or service; each server sees only its own card
CUDA_VISIBLE_DEVICES=0 llama-server -m Meta-Llama-3.1-8B-Instruct-Q8_0.gguf -c 16384 -ngl 99 --port 8080
CUDA_VISIBLE_DEVICES=1 llama-server -m Meta-Llama-3.1-8B-Instruct-Q8_0.gguf -c 16384 -ngl 99 --port 8081
```

vLLM can run both copies behind one address with `--data-parallel-size 2`. The model still has to fit on one card, and an 8B model at 16 bits is about 16 GB (calculated: 8.03 billion parameters × 2 B). So pick a 4-bit build, for example Qwen2.5-7B's official AWQ build at 5.2 GiB (published file size on Hugging Face):

```
vllm serve Qwen/Qwen2.5-7B-Instruct-AWQ --data-parallel-size 2 --max-model-len 16384
```

**Speed ceiling (estimate):** 8.54 GB ÷ 360 GB/s ≈ 24 ms per token, so at most ≈ 42 tokens/s per stream, with two streams running at once. Two copies of the same model let more users share it. Two different models let you run two services, for example a chat model and a coding model.

### Which one

    |  | A: one 32B model, split | B: two 8B models, one per card |  |

    | Model | One larger model (32B, 4-bit) | Two smaller models (8B, 8-bit) |  |

    | Single-stream ceiling (estimate) | ≈ 18 tokens/s with a layer split, ≈ 36 with a tensor split | ≈ 42 tokens/s each |  |

    | Context per request (estimates above) | 8k comfortable, ≈ 14k at the edge | 16k each, with ≈ 1 GiB left |  |

    | Pick it when | The task needs the larger model's quality | You serve many requests, or two different jobs |  |

## The common mistake: adding the cards up

"24 GB in total, so a 23 GB model fits." Here is what actually happens with Qwen2.5-32B at Q5_K_M:

  - The file is 23,262,157,696 bytes, which is 21.7 GiB (published file size on Hugging Face).

  - Split in two, each card holds 10.8 GiB of weights plus ≈ 0.75 GiB of overhead (estimate), out of the ≈ 11.75 GiB CUDA can use.

  - That leaves ≈ 0.2 GiB for the KV cache per card, which is about 1,400 tokens at 128 KiB per token (estimate).

Depending on the runtime and its settings, one of three things happens:

  - the load fails;

  - the context is cut down to what fits;

  - some layers stay in system RAM, where generation is much slower.

Check the startup log to see which one you got. llama.cpp prints how many layers it put on the GPUs (`offloaded 65/65 layers to GPU`: 64 layers plus the output layer). vLLM prints the KV cache it got (`GPU KV cache size: … tokens`).

The same habit of treating two cards as one bigger card causes two related mistakes:

  - expecting one request to run twice as fast;

  - expecting every tool to use both cards.

## On AxForge

  - RTX 3060: rent one card or both cards, and try examples A and B on the two-card machine.

  - GPU pages: memory, bandwidth and labelled community figures for each GPU.

  - DGX Spark: unified memory on one device, for when you want the room without a split.

  - How GPU rentals work: SSH, the tools on the machine, and how the clock runs.

  - Language models: find a build that fits your cards.

  - Quickstart: or skip the hardware and call Qwen3.8 27B through the API.

## Next

  - Model size, GPU memory and CPU offload: the fit arithmetic for a single card.

  - Context windows and the KV cache: where the per-token numbers come from.

  - llama.cpp vs vLLM: one stream vs many, on one card or several.



Source: https://dev.axforge.ai/ai-engineering/multi-gpu/
