How AI inference works
Your application sends text through a client to an inference server, which turns it into tokens and runs a model's weights on one or more GPUs: one pass over the whole input (prefill), then one output token at a time (decode) until it stops, counting the tokens both ways.
Five layers, five jobs
Every AI call passes through the same layers, whether the model is hosted or you run it yourself. Most confusion comes from treating two of these layers as one.
- Application
- Your product: a web app, a CLI, an IDE plug-in, a nightly batch job. It decides what to ask and what to do with the answer.
- Agent, SDK or client
- The code that talks to the server. An SDK (the OpenAI SDK, for example) turns a function call into an HTTP request. An agent harness (aider, Cline, Claude Code, Codex CLI) adds a loop: call the model, run the tools it asks for (read a file, run the tests), send the results back, then call again. Each turn of that loop is a separate inference request. The client also holds the conversation. A chat completions API keeps no state between requests, so the full history goes with every request.
- Inference server (runtime)
- The program that loads a model and serves it: vLLM, llama.cpp's
llama-server, SGLang, TensorRT-LLM, Ollama and others. It applies the chat template and tokenizes the input. It queues and batches requests, manages the KV cache and runs the model. Then it samples tokens, streams them back and counts usage. Runtimes are built for different workloads. vLLM is built for many requests at once (continuous batching, paged KV cache).llama-serveris built for GGUF files, simple setup and running on CPU and GPU together, and it also serves several requests at once in parallel slots. Neither runtime keeps a GPU busy on its own: during decode, one stream leaves most of the GPU's compute idle in any runtime, and only enough requests at once fill it. See llama.cpp vs vLLM. - Model
- Files, not a program. A model is its weights (billions of numbers learned in training), a config (layers, attention heads, context length), a tokenizer and a chat template. Qwen3.8 27B has about 27 billion weights (vendor figure). One model comes in several builds (BF16, FP8, 4-bit). Smaller builds use less memory; how much quality they lose depends on the quantization method. See Quantization. A model does nothing until a runtime loads it.
- GPU(s)
- The hardware that does the math. Its memory holds the weights and the KV cache, its cores run the matrix multiplications, and its memory bandwidth sets how fast the weights can be read. A model fits only if the weights, the KV cache, the runtime's overhead and its working memory all fit together. If they do not, some runtimes keep part of the model in system RAM and run those layers on the CPU. That works, but it is much slower (see Model fit). Two GPUs are two separate memories. A runtime that supports it can split one model across them, but data then crosses PCIe, and one request does not get twice as fast (see One vs multiple GPUs). Some machines have unified memory instead: the NVIDIA DGX Spark (GB10) has 128 GB (vendor figure) that the CPU and GPU share.
Tokens, context and the two phases
- Prompt (input)
- Everything you send: the system message, the conversation so far, tool results, pasted files. The runtime puts all of it through the model's chat template to make one sequence, which adds role markers you never typed.
- Tokens
- The units a model reads and writes: words, word pieces, punctuation, runs of whitespace. English prose averages roughly 4 characters per token on many tokenizers (estimate). Code, numbers and other languages split differently. Each model has its own tokenizer, so the same text gives a different count on a different model. APIs measure usage in tokens.
- Context
- All the tokens the model sees in one request: the input plus the output written so far. The context window is the most it can take. Qwen3.8 27B accepts 262,144 tokens (vendor figure: the model's published context length). That is an upper limit, not a recommended size. Every token of context costs KV-cache memory and prefill time. See Context and the KV cache.
- Inference
- Running a trained model to get an output. Training is a different job and is not part of an API call.
Inference has two phases. Prefill reads the whole input in one pass. All input tokens go through the model together, layer by layer; the runtime stores each token's keys and values for each layer (the KV cache), and the last position produces the first output token. Prefill is mostly limited by compute, takes longer as the input grows, and decides how long you wait for the first token. Decode then writes the reply one token at a time. Each step runs only the newest token through the model, reads the weights and the KV cache from memory, and samples the next token. It continues until the model emits an end-of-sequence token, hits a stop sequence or reaches max_tokens. For a single request, decode speed is limited mostly by memory bandwidth, not by compute. That is why servers batch the decode steps of many requests together: one read of the weights serves every sequence in the batch. Batching raises total throughput (tokens per second across all users). It does not make any one stream faster, and each stream usually gets a little slower.
Not every endpoint generates text. An embedding model such as Qwen3 Embedding runs prefill over the input and returns one vector, with no token-by-token output. A decision endpoint such as Xev (POST /v1/systemone) returns a yes/no, choice or score with probabilities instead of free text.
Serverless API or a dedicated machine
It is the same stack either way. What changes is which layers you run and which the provider runs.
| Serverless API | Dedicated machine | |
|---|---|---|
| Who runs the server | The provider | You, over SSH |
| What you choose | A model from the served list, and request parameters | Model, build, runtime and every runtime setting |
| What you pay for | Usually tokens in and out | Time the machine is yours, busy or idle |
| Capacity | Shared with other users, within rate limits | The whole GPU, and only that GPU |
| Setup | An API key and a base URL | Download weights, start a runtime, reach its port |
| Fits when | A served model does the job; load is bursty or unknown | You need a model or setting that is not served, or steady load keeps the GPU busy |
Both usually accept the same OpenAI-compatible requests. On a dedicated machine you start the runtime yourself. Either of these serves /v1/chat/completions on the machine:
# vLLM: Hugging Face weights, serves on port 8000
vllm serve Qwen/Qwen2.5-7B-Instruct --max-model-len 32768
# llama.cpp: one GGUF file, 8,192-token context, all layers on the GPU, serves on port 8080
llama-server -m ./model-Q4_K_M.gguf -c 8192 -ngl 99
The vLLM command loads BF16 weights of about 15.2 GB (calculated: 7.61 billion parameters, vendor figure, × 2 bytes). It also needs KV cache for at least one full 32,768-token sequence: 2 (keys and values) × 28 layers × 4 KV heads × 128 head_dim × 2 bytes = 57,344 bytes per token, × 32,768 tokens ≈ 1.9 GB (calculated from the model's published config). That is over 17 GB before runtime overhead, so it does not fit a 12 GB card; a 4-bit GGUF of the same model does (see Model fit). To use either one, point your client at http://localhost:8000/v1 or http://localhost:8080/v1 instead of a hosted base URL. When the machine is remote, use an SSH tunnel (llama-server listens only on 127.0.0.1 unless you set --host). The application code stays the same.
The stack in one diagram
APPLICATION web app, CLI, IDE plug-in, batch job
|
v
AGENT / SDK / CLIENT builds the messages, keeps the history,
| runs tools, retries
|
| HTTPS POST /v1/chat/completions
| JSON in; JSON or a token stream (SSE) out
- - - -|- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
| above: your code below: the provider's (serverless)
v or yours (dedicated machine)
INFERENCE SERVER auth, queue, chat template, tokenizer,
(runtime) batching, KV cache, sampling, streaming,
| usage counts
v
MODEL weights + config + tokenizer + chat template
| (files on disk, loaded into GPU memory)
v
GPU(s) memory: weights + KV cache
cores: the matrix math
Worked example: one chat request
A user types Explain a hash map in two sentences. into a chat feature. This is what happens to that request, step by step.
- The client sends. The SDK posts the model id, the messages (a system message, the earlier turns, the new question),
max_tokensandstream: true. The whole conversation is sent every time. - The gateway routes. A hosted API checks the key and passes the request to a server that has the model loaded. If that server is busy, the request waits in its queue.
- Template and tokenizer. The messages become one sequence with role markers, then token ids. The question alone is about 8 tokens; with the template's role markers it is about 20, and with a system message and earlier turns it is hundreds or thousands (estimates).
- Prefill. All input tokens go through the model in one pass. The KV cache now holds an entry per token per layer. The model outputs a probability for every token in its vocabulary. The sampler (temperature, top_p) picks the first output token, and it is streamed back.
- Decode. That token goes back in as input. Each step reads the weights, attends over the KV cache, adds one KV entry, samples the next token and streams it. This repeats for every token of the reply.
- Stop. The model emits its end-of-sequence token (
finish_reason: "stop") or the reply hitsmax_tokens(finish_reason: "length"). The runtime then frees the memory that held this request's KV cache. A runtime with prefix caching may keep it, so the next request that starts with the same text (such as the next turn of this conversation) can skip that part of prefill. - Usage. The last chunk reports
prompt_tokens,completion_tokensandtotal_tokens.prompt_tokenscounts the input after the template, so it is more than the words you typed.completion_tokenscounts everything generated, reasoning included. Many servers (OpenAI's API and vLLM among them) put usage in a stream only when the request setsstream_options: {"include_usage": true}, so the code below sets it.
request sent
|--network--|--queue--|------ prefill ------|-d-|-d-|-d-|-d-| ... |-d-| stop, usage
^ ^
first token last token
time to first token = network + queue + prefill grows with INPUT length
time after that = output tokens x one decode step grows with OUTPUT length
For the arithmetic, assume a decode speed of 30 tokens/s (estimate, for illustration only). A 60-token answer then takes 60 / 30 = 2 s of decode; a 1,500-token answer, or a long reasoning trace before a short answer, takes 1,500 / 30 = 50 s (both calculated). Pasting a long file into the prompt mostly delays the first token. The steps after it slow down only a little, because each one also reads a larger KV cache.
To see the two phases with your own numbers, time the first token and the rest separately:
# pip install openai · the key is read from AXFORGE_API_KEY, never written here
import os, time
from openai import OpenAI
client = OpenAI(base_url="https://api.axforge.ai/v1", api_key=os.environ["AXFORGE_API_KEY"])
t0 = time.perf_counter()
first = usage = None
stream = client.chat.completions.create(
model="qwen3.8-27b-nvfp4",
max_tokens=200,
stream=True,
stream_options={"include_usage": True}, # usage in the last chunk
messages=[{"role": "user", "content": "Explain a hash map in two sentences."}],
extra_body={"chat_template_kwargs": {"enable_thinking": False}}, # the answer only, no reasoning trace
)
for chunk in stream:
if first is None and chunk.choices and chunk.choices[0].delta.content:
first = time.perf_counter()
if chunk.usage:
usage = chunk.usage
secs = time.perf_counter() - first
rest = usage.completion_tokens - 1 # the tokens after the first one
print(f"first token after {first - t0:.2f} s (network + queue + prefill)")
print(f"decode: {rest} more tokens in {secs:.2f} s, about {rest / secs:.0f} tokens/s")
print(f"prompt: {usage.prompt_tokens} tokens")
// npm i openai · save as journey.mjs, run: node journey.mjs
// the key is read from AXFORGE_API_KEY, never written here
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.axforge.ai/v1", apiKey: process.env.AXFORGE_API_KEY });
const t0 = performance.now();
let first, usage;
const stream = await client.chat.completions.create({
model: "qwen3.8-27b-nvfp4",
max_tokens: 200,
stream: true,
stream_options: { include_usage: true }, // usage in the last chunk
messages: [{ role: "user", content: "Explain a hash map in two sentences." }],
chat_template_kwargs: { enable_thinking: false }, // the answer only, no reasoning trace
});
for await (const chunk of stream) {
if (first === undefined && chunk.choices[0]?.delta?.content) first = performance.now();
if (chunk.usage) usage = chunk.usage;
}
const secs = (performance.now() - first) / 1000;
const rest = usage.completion_tokens - 1; // the tokens after the first one
console.log(`first token after ${((first - t0) / 1000).toFixed(2)} s (network + queue + prefill)`);
console.log(`decode: ${rest} more tokens in ${secs.toFixed(2)} s, about ${(rest / secs).toFixed(0)} tokens/s`);
console.log(`prompt: ${usage.prompt_tokens} tokens`);
Run it twice: once as it is, then with a few thousand words pasted into the question. The first line grows with the input. The decode speed in tokens per second changes much less.
The common mistake
"The model is the API." An API is a contract: a URL, a JSON format and a key. Behind it are a runtime, a particular build of the model, and settings. The same model name on two endpoints can answer differently. The build may differ (BF16 or 4-bit), and so may the chat template, the maximum context the server allows, the default sampling settings, and how tool calls and reasoning are parsed. Changing the base URL can change all of these, not just the address. When results differ between providers, compare these first before blaming the model.
"A bigger GPU makes everything faster." Each property of a machine buys something different:
| More of | Buys | Does not buy |
|---|---|---|
| GPU memory | Bigger models, longer contexts, more requests at once | Faster tokens for one request |
| Memory bandwidth | Faster decode for one request | Room for a bigger model |
| Compute | Faster prefill; more throughput when many requests are batched | Much faster decode for one request |
| GPUs | Room for one model split across cards, or more separate streams | One request twice as fast |
Here is a worked comparison for one stream of decode on a dense model. Each output token reads every weight once, so memory bandwidth divided by the size of the weights gives a rough ceiling for plain decoding (calculated). Real numbers are lower, and techniques such as speculative decoding can change the picture. Take an 8B model at 4-bit, which is about 4.9 GB of weights as a Q4_K_M GGUF file (estimate):
| RTX 3060 12 GB | DGX Spark (GB10) | |
|---|---|---|
| Memory | 12 GB GDDR6 (vendor figure) | 128 GB unified (vendor figure) |
| Memory bandwidth | 360 GB/s (calculated: 15 Gbps × 192-bit bus ÷ 8) | 273 GB/s (vendor figure) |
| Decode ceiling, one stream | 360 / 4.9 ≈ 73 tokens/s (calculated) | 273 / 4.9 ≈ 56 tokens/s (calculated) |
The machine with more than ten times the memory (128 / 12 ≈ 10.7, calculated) has the lower ceiling for this one stream. Its advantage is the rest of the first table: models and contexts that never fit in 12 GB, many more requests at once, and faster prefill. Choose hardware for the workload (model size, context, how many requests at once), not for one headline number.
On AxForge
- Send your first request: Quickstart, then Chat completions for streaming, reasoning output and tool calls.
- The served models and their ids: Models. Every model, served or not: the catalogue.
- Inference that does not write text: Embeddings and Xev decisions.
- An agent or editor on the API: Use it from your stack.
- A dedicated machine: How GPU rentals work, RTX 3060, DGX Spark.
- What tokens and machine time cost: Pricing. Keys and usage: the console.
Next
- Model size, GPU memory and CPU offload: will this model fit, and why is it using the CPU?
- Context windows and the KV cache: what context costs in memory.
- llama.cpp vs vLLM: one stream or many.