AI Engineering / Models & Memory

Context windows and the KV cache

The context window is the most tokens a model can attend to in one sequence; the KV cache is the memory those tokens occupy while a request runs, and because it grows with every token and every concurrent sequence, at long contexts it can outweigh the model itself.

Context window is not model size

Every model card gives both numbers, and they measure different things.

Model sizeContext window
What it countsWeights (parameters)Tokens one sequence can hold: prompt plus output
Set byThe architectureTraining and position-encoding scaling
MemoryFixed, loaded once, shared by every requestGrows per token, per sequence, through the KV cache
Llama 3.1 8B8.03 B parameters, about 16 GB at BF16131,072 tokens

Parameter count and window: vendor figures (Meta's published model and config.json). Weight size: calculated, 8.03 B × 2 bytes.

Model size decides whether the weights fit. The context window is a ceiling: something the model can do, not a setting to max out. What a request costs in memory is the context it uses, or, on runtimes that reserve memory up front, the context you configure. That cost lives in the KV cache.

What the KV cache stores

To produce each new token, every attention layer compares it with every earlier token in the sequence. It does this through two vectors per earlier token, per layer: a key (K), used to decide how relevant that token is, and a value (V), the information taken from it. Recomputing them for the whole history at every step would waste most of the work, so the runtime computes them once and keeps them. That store is the KV cache.

one sequence, one layer, and the same again in every layer

token    The   invoice   is   due   on   Friday   → next token
key      K1    K2        K3   K4    K5   K6
value    V1    V2        V3   V4    V5   V6
         └─── kept until the request ends ────┘

the next token is scored against every K and mixes every V
  • Prompt and output both count. A 6,000-token prompt that produces 2,000 tokens ends with 8,000 tokens in the cache.
  • It belongs to one sequence. Eight requests running together hold eight caches. A runtime can share the cache of an identical prefix, such as a common system prompt (prefix caching), but the rest of each request is cached separately.
  • It is read in full for every new token. As the sequence grows, each output token moves more bytes through memory, so generation slows as a conversation gets longer.

How big it gets

For standard attention, the size per token follows from four numbers in the model's published config (config.json on Hugging Face):

bytes per token = 2 × layers × kv_heads × head_dim × bytes_per_value
                  │   │        │          │          └─ 2 at FP16/BF16, 1 at FP8
                  │   │        │          └─ width of one attention head
                  │   │        └─ key/value heads (fewer than query heads under GQA)
                  │   └─ attention layers whose cache grows
                  └─ one K and one V

KV cache = bytes per token × tokens per sequence × sequences running at once

The second formula is the one to remember: the cache grows linearly with context and linearly with concurrency. Doubling either one doubles the memory. For Llama 3.1 8B the cache costs 128 KiB per token (worked out below). For a single sequence, that gives:

KV cache for ONE sequence · Llama 3.1 8B · FP16 cache · 128 KiB per token
calculated from the model's published config · each █ = 1 GiB · 8K = 8,192 tokens

   8K  █                                   1 GiB
  32K  ████                                4 GiB
 128K  ████████████████                   16 GiB   this model's limit
 256K  ████████████████████████████████   32 GiB   past the limit, for scale

       ███████████████                   ~15 GiB   the weights at BF16, for comparison

At its full 131,072-token window, one sequence of this 8B model needs more memory for its cache than for its weights.

The attention design sets the cost per token

The cache per token depends on how a model does attention, not on its parameter count. Every figure in this table is at 16 bits per value, calculated from the model's published config with the formula above.

ModelLayers whose cache growsKV headshead_dimKV per token
Llama 2 7B (no GQA)3232128512 KiB
Llama 3.1 8B328128128 KiB
Qwen2.5-7B28412856 KiB
Gemma 3 27B (sliding window)10 of 621612880 KiB + up to 416 MiB per sequence
Qwen3-Next-80B-A3B (hybrid)12 of 48225624 KiB + a fixed state per sequence
  • Grouped-query attention (GQA) lets several query heads share one K/V head. Llama 3.1 8B has 32 query heads but only 8 KV heads, so its cache is a quarter of what full multi-head attention would need. Compare Llama 2 7B, which has no GQA.
  • Sliding-window attention limits a layer to the last N tokens, so that layer's cache stops growing at N. Gemma 3 27B uses a 1,024-token window in five of every six layers (Google's published config), so only 10 of its 62 layers grow with context: 80 KiB per token, against 496 KiB if all 62 kept full attention, plus 52 windowed layers × 1,024 tokens × 8 KiB = 416 MiB once a sequence passes 1,024 tokens (calculated). The runtime has to support this for the model. If it does not, it allocates a full cache anyway.
  • Hybrid and linear attention replace many attention layers with a recurrent state of fixed size, either linear attention such as Gated DeltaNet or state-space layers such as Mamba. In Qwen3-Next-80B-A3B, 36 of the 48 layers keep a fixed state of about 19 million values per sequence (36 layers × 32 heads × 128 × 128, about 36 MiB at 16 bits, plus a small convolution state; calculated). That state is the same size at 1,000 tokens or 200,000. Only the 12 full-attention layers grow per token, so this 80B model needs less cache per token than an 8B one.
  • Multi-head latent attention (DeepSeek-V3) stores one compressed vector per token per layer instead of full K and V, so the formula above does not apply as written. From the published config: 61 layers × (512 + 64) values × 2 bytes = 70,272 bytes, about 69 KiB per token for a 671 B-parameter model (calculated).

A 1M-token window is not a 1M-token request

Some models publish windows of a million tokens or more. That number says what the model can accept, not what a request should carry.

  • Memory. A model shaped like Llama 3.1 8B would need 128 KiB × 1,048,576 tokens = 128 GiB of FP16 cache for one such sequence (calculated; the real model stops at 131,072). Efficient attention designs make this smaller, but the cost does not go away.
  • Time. The whole prompt is processed before the first output token appears (prefill). In full-attention layers, the attention part of that work grows with the square of the prompt length, so time to first token rises steeply.
  • Money. Hosted APIs bill input tokens, and a chat or agent resends its history on every turn.
  • Quality. Published research found that models use information from the middle of a long prompt less reliably than from its start or end (Liu et al., “Lost in the Middle”, 2023). The RULER benchmark (Hsieh et al., 2024) found that for many models the effective context is shorter than the advertised window.

Send what the task needs: retrieve the relevant passages with embeddings, summarise old turns and trim tool output. Keep the long window for the rare request that really needs it.

Worked example: Llama 3.1 8B

The published config gives 32 layers, 8 key/value heads and a head_dim of 128 (hidden size 4,096 ÷ 32 query heads). Every number in the steps and the table is calculated from those values.

  1. Per token, FP16 cache: 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB.
  2. One 8K sequence: 131,072 bytes × 8,192 tokens = 1,073,741,824 bytes = 1 GiB.
  3. One 128K sequence: 131,072 bytes × 131,072 tokens = 17,179,869,184 bytes = 16 GiB.
  4. Eight concurrent sequences: 8 × 1 GiB = 8 GiB at 8K; 8 × 16 GiB = 128 GiB at 128K.
  5. FP8 cache: 1 byte per value instead of 2 halves every figure.
Llama 3.1 8B, calculatedFP16 cacheFP8 cache
Per token128 KiB64 KiB
One 8K sequence1 GiB0.5 GiB
One 128K sequence16 GiB8 GiB
8 sequences × 8K8 GiB4 GiB
8 sequences × 128K128 GiB64 GiB
Weights at BF16 (8.03 B × 2 bytes), for comparisonabout 16 GB (15 GiB)

The cache also affects speed. At 128K, each output token of one sequence reads 16 GiB of cache plus about 15 GiB of weights (calculated), about twice the bytes of a short request. On memory-bound hardware that roughly halves single-stream generation speed (estimate).

How that plays out on real hardware: a model fits only if weights, KV cache, runtime overhead and working memory all fit at once.

  • One 12 GB card (12,288 MiB, vendor figure) running the Q4_K_M GGUF of this model (4.92 GB, the file size published on Hugging Face) with the default FP16 cache: about 7.4 GiB is left after the weights (calculated), or about 6 GiB after runtime overhead (estimate). At 128 KiB per token, that holds about 49,000 tokens of cache in total (calculated): six 8K conversations, or one 32K conversation with room to spare. A single 128K request does not fit.
  • A 128 GB unified-memory machine (vendor figure) with BF16 weights: eight 128K sequences need 128 GiB (137 GB) of FP16 cache, more than the whole machine before the weights are even loaded. With an FP8 cache it is 64 GiB plus about 15 GiB of weights, roughly 79 GiB or 85 GB (calculated). That fits on paper and leaves room for the operating system and the runtime (estimate).

Making the cache smaller

  • Cap the context at the longest request you actually send (prompt plus output). This is the cheapest saving, and it costs no quality.
  • Cap concurrency when long requests must fit: --max-num-seqs in vLLM, -np in llama.cpp.
  • Quantize the cache. An FP8 cache is half the size of FP16. llama.cpp's q8_0 stores 32 values in 34 bytes, about 53 % of FP16, and q4_0 stores them in 18 bytes, about 28 % (calculated from the block formats). At 8 bits the quality change is usually small (community measurements), but it depends on the model and the method, so test it on your own prompts. Whether a format is available depends on the runtime version, the GPU and the attention backend.
  • Pick the architecture for the job. If long context is the workload, a GQA, sliding-window or hybrid model needs a fraction of the cache per token.

vLLM with a context cap and an FP8 cache:

vllm serve meta-llama/Llama-3.1-8B-Instruct --max-model-len 16384 --kv-cache-dtype fp8

llama.cpp with a 32K cache at 8 bits. A quantized V cache needs flash attention; current builds switch it on by themselves, and -fa on makes that explicit:

llama-server -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -c 32768 -ngl 99 -fa on -ctk q8_0 -ctv q8_0

In current llama-server, -c is one cache for the whole server. With the default slot count (automatic: four slots sharing one buffer), a single request can use all of it; with an explicit -np N, each slot gets -c ÷ N unless --kv-unified is set. This has changed between releases, so check n_ctx_seq in the startup log. vLLM reports its pool at startup too: the KV cache size in tokens, and the maximum concurrency at your --max-model-len.

The common mistake: maximum context “just in case”

The model supports 128K, so someone starts the server with 128K. What happens next depends on how the runtime holds the cache:

How the runtime holds the cacheWhat “maximum, just in case” does
Allocated up front for the configured context (llama.cpp -c, Ollama num_ctx)The memory is taken at load time, used or not. Either the model fails to load or, where the runtime places layers itself (llama.cpp's automatic -ngl, Ollama), fewer layers stay on the GPU and the rest run on the CPU. That makes inference much slower.
A paged pool sized from the memory left after the weights (vLLM)--max-model-len defaults to the model's own maximum, and the server refuses to start if one sequence that long does not fit. --max-model-len auto picks the longest that fits: the same choice, made for you. Once running, a few long requests can fill the pool while the others wait or are preempted.
Grows as tokens arrive (a plain PyTorch or Transformers loop)Short test prompts work. The first long prompt in production runs out of memory mid-request.

The fix is to size from the workload, not from the model card:

  1. Measure the longest prompt plus output you actually send, and set the context a little above that.
  2. Spend the memory this frees on concurrency, and confirm the result in the startup log.
  3. If a few requests really need a very long context, give them their own deployment instead of sizing every request for them.

On AxForge

  • Qwen3.8 27B on the API accepts up to 262,144 tokens. It uses hybrid attention, so its cache grows more slowly per token than a full-attention model of its size. Time to first token and the input tokens you are billed for still grow with every token you send. Model ids and limits are in Models; the request shape and max_tokens are in Chat completions.
  • To send less, embed and retrieve: Embeddings with Qwen3 Embedding.
  • Coding agents resend the conversation and files on every turn. Tell the tool the model's context window so it trims history before the limit: aider, Cline.
  • To size a cache yourself, rent a machine and run your own runtime: How GPU rentals work, RTX 3060 (12 GB per card), DGX Spark (128 GB unified memory).
  • Each model's published context window is in the model catalogue. Token costs are on the pricing page.

Next

Ask on the forum Markdown For AI Updated 2026-10-01

Anything unclear on this page?

Ask on the forum — the answer helps the next person too.

Ask about this page